Method and device for determining a search rearrangement model
The search re-ranking model optimizes search result ordering by training an initial scoring model with an evaluation model to adjust scores based on expected rankings, enhancing accuracy and user experience.
Patent Information
- Application Number
- CN202210367936.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-08
AI Technical Summary
In the existing search system, the accuracy of search results sorting is insufficient, resulting in inefficient users' access to information.
The initial sorting and scoring of search result entries is performed through the initial scoring and sorting model. The scoring evaluation model is used to determine the reward score based on the expected sorting, calculate the loss function and perform model training, and obtain the target scoring and sorting model to optimize the sorting results.
It improves the accuracy of search results sorting and improves users' search experience and information acquisition efficiency.
Smart Images

Figure CN114722086B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method and apparatus for determining a search rearrangement model. Background Art
[0002] With the rapid development of information technology, online search has become one of the important ways for people to obtain information. Specifically, after a user enters a search term in a search system, the system will recall a large number of search result entries, and perform initial sorting and refined sorting on them, and finally display some of them to the user.
[0003] In a search system, the quality of search sorting greatly affects the page quality. Specifically, it refers to the degree of match between the search results displayed on the page presenting search results to the user and the user's search expectation. When the page quality is poor, it will present a large amount of redundant information to the user, reducing the efficiency of the user to obtain information; when the page quality is good, it will display a page that is more in line with the user's expected content for the user, improving the efficiency of the user to obtain information.
[0004] It can be seen that how to improve the accuracy of the sorting result of the search result entries to be sorted is of great significance for improving the user's search experience. Summary of the Invention
[0005] To solve the above technical problems, this application provides a method and apparatus for determining a search rearrangement model, which improves the accuracy of sorting the search results for a target search term.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On the one hand, the embodiments of this application provide a method for determining a search rearrangement model, and the method includes:
[0008] Obtain a training sample set corresponding to a target search term; the training sample set includes the target search term and multiple search result entries of the target search term;
[0009] Input the training sample set into an initial scoring and sorting model, and determine the initial sorting scores of the multiple search result entries respectively through the initial scoring and sorting model;
[0010] Input the training sample set including target sorting labels and the initial sorting scores into a scoring and evaluation model, and determine the reward scores of the initial sorting scores respectively through the scoring and evaluation model based on the target sorting labels and the initial sorting scores; wherein, the target sorting labels are used to identify the expected sorting of the multiple search result entries in the training sample set;
[0011] Determine a loss function for the multiple search result entries based on the initial sorting score and the reward score;
[0012] Perform sorting model training on the initial scoring and sorting model according to the loss function to obtain a target scoring and sorting model for performing search rearrangement on the multiple search result entries.
[0013] On the other hand, an embodiment of the present application provides a device for determining a search rearrangement model. The device includes an acquisition unit, a determination unit, and a training unit:
[0014] The acquisition unit is configured to acquire a training sample set corresponding to a target search term; the training sample set includes the target search term and multiple search result entries of the target search term;
[0015] The determination unit is configured to input the training sample set into an initial scoring and sorting model, and determine respective initial sorting scores of the multiple search result entries through the initial scoring and sorting model;
[0016] The determination unit is further configured to input the training sample set including the target sorting label and the initial sorting score into a scoring evaluation model, and determine respective reward scores of the initial sorting scores through the scoring evaluation model based on the target sorting label and the initial sorting score; wherein, the target sorting label is used to identify the expected sorting of the multiple search result entries in the training sample set;
[0017] The determination unit is further configured to determine a loss function for the multiple search result entries based on the initial sorting score and the reward score;
[0018] The training unit is configured to perform sorting model training on the initial scoring and sorting model according to the loss function to obtain a target scoring and sorting model for performing search rearrangement on the multiple search result entries.
[0019] As can be seen from the above technical solution, the initial sorting score model is used to perform initial sorting and scoring on multiple search result entries of the target search term to obtain their respective corresponding initial sorting scores. The scoring evaluation model evaluates the initial sorting scores output by the initial sorting score model according to the expected sorting of the multiple search result entries, and determines their respective reward scores. Further, the loss function of the multiple search result entries is determined according to the initial sorting scores and the reward scores, and the initial sorting score model is trained according to the loss function to obtain a target sorting score model for performing search rearrangement on the multiple search result entries. It can be seen that by using the scoring evaluation model to determine the reward scores according to the expected sorting for the initial sorting scores output by the initial sorting score model, the purpose of unsupervised training of the initial sorting score model is achieved, and the reward scores are determined based on the expected sorting. Therefore, by using the loss function determined according to the initial sorting scores and the reward scores to train the initial sorting score model for the sorting model, the initial sorting score model can be optimized in the direction of outputting the expected sorting, and finally the target sorting score model is obtained, improving the accuracy of the sorting result. The target sorting score model is used to perform search rearrangement on the multiple search result entries, so as to display the relevant search results of the target search term for the user according to the sorting result after the search rearrangement. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a flowchart of a method for determining a search rearrangement model provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic framework diagram of a method for determining a search rearrangement model provided by an embodiment of the present application;
[0023] Figure 3 It is a device structure diagram of a device for determining a search rearrangement model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of this application.
[0025] Figure 1 It is a flowchart of a method for determining a search rearrangement model provided by an embodiment of this application. The method includes S101 - S105:
[0026] S101: Obtain a training sample set corresponding to the target search term.
[0027] Among them, the training sample set includes the target search term and multiple search result entries of the target search term. Specifically, the training sample set can be constructed in the following way:
[0028] According to a series of search result entries obtained from a single search of the target search term by the user, a certain number of them are selected as a single List sample in the training sample set. For example, for the target search term query A, M search result entries are obtained from a single search, and N search result entries and query A are jointly used as the List sample of query A, where N ≤ M.
[0029] To prevent dirty data from affecting the accurate learning of other data in the training sample set, during the construction of the training sample set, truncation of the List length can be performed. Specifically, for a series of target search terms query and their corresponding search result entries, the search term query with the number of search result entries less than the preset List length is discarded. For example, if the aforementioned N is 10 and the number of search result entries of query B is 5, then this group of data of query B is discarded. It should be noted that for the preset List length, it can be set according to the actual training sample requirements, and this application does not make any limitations on this. In addition, the training sample set can include multiple List samples.
[0030] S102: Input the training sample set into the initial scoring and ranking model, and determine the initial ranking scores of the multiple search result entries through the initial scoring and ranking model.
[0031] The initial scoring and ranking model is used to score multiple search result entries of the target search term in the training sample set, obtain the initial ranking scores of the multiple search result entries, and based on this, determine the initial ranking of the multiple search result entries of the target search term.
[0032] For the input method of inputting the training sample set into the initial scoring and ranking model, multiple List samples in the training sample set can be successively input into the initial scoring and ranking model in the manner of a single List sample, or multiple List samples in the training sample set can be sorted according to the difficulty of learning first, and then successively input into the initial scoring and ranking model in the order from easy to difficult, or it can be set according to actual training requirements. The present application does not make any limitation on this.
[0033] In a possible implementation manner, S102 includes the following steps:
[0034] S1021: Input the training sample set into the initial scoring and ranking model, and determine the feature vectors of multiple search result entries in the training sample set through the feature extraction layer of the initial scoring and ranking model;
[0035] S1022: Determine the initial ranking scores of the multiple search result entries according to the feature vectors.
[0036] In a possible implementation manner, the feature vector includes a basic feature vector and a key statistical feature vector.
[0037] Specifically, after inputting the training sample set into the initial scoring and ranking model, the basic features including the refined ranking scores, relevance scores, cold start scores, policy rule scores, etc. of multiple search result entries including the target search term are extracted through the feature extraction layer of the initial scoring and ranking model and vectorized to obtain the basic feature vector; the key statistical features including the historical click-through rate, historical browsing duration or playing duration of the search result entries are extracted and vectorized to obtain the key statistical feature vector; based on this, the context feature including the basic feature vector and the key statistical feature vector is obtained. In addition, the ranking positions of multiple search result entries can be extracted to obtain the position vector, which is connected to the ranking model training layer, thereby enriching the context features of the search result entries.
[0038] It should be noted that for the categories of the basic vector and the key statistical feature vector, they can be set according to specific actual training requirements, and the present application does not make any limitation on this. For example: when the target search term and its search result entries are of video type, the historical click-through rate and historical playing duration can be selected as the key statistical features; when the target search term and its search result entries are of text type, the historical click-through rate and historical browsing duration can be selected as the key statistical features, thereby enabling the finally trained target ranking and scoring model to better meet the requirements of specific usage scenarios.
[0039] S103: Input the training sample set including the target sorting label and the initial sorting scores into the scoring and evaluation model. Based on the target sorting label and the initial sorting scores, the scoring and evaluation model determines the reward scores for the respective initial sorting scores.
[0040] Among them, the target sorting label is used to identify the expected sorting of the multiple search result entries in the training sample set, and is used to represent the optimal sorting result of the multiple search result entries for the current target search term. Specifically, the key statistical indicators that are most concerned in the specific usage scenario can be used as the basis for target sorting, and based on this, the target sorting labels of the multiple search result entries for the target search term are determined to identify the expected sorting of the multiple search result entries.
[0041] The scoring and evaluation model is used to evaluate the initial sorting scores output by the initial scoring and sorting model. Specifically, it determines the reward scores for the initial sorting scores of the multiple search result entries according to the target sorting label.
[0042] In order to make the evaluation of the initial sorting scores output by the initial scoring and sorting model by the scoring and evaluation model more accurate, in a possible implementation, the time difference method is used to train the evaluation model of the scoring and evaluation model. That is, the evaluation model of the scoring and evaluation model is trained in the way of TD difference, and the training efficiency of this method is relatively high compared with other methods.
[0043] In a possible implementation, the scoring and evaluation model can be used to match the initial sorting of the multiple search result entries with the expected sorting of the multiple search result entries to determine the reward scores for the respective initial sorting scores; the initial sorting of the multiple search result entries is determined based on the initial sorting scores. For example, for the List sample of query A, the expected sorting of N search result entries is determined according to the target sorting label, the initial sorting of N search result entries is determined according to the initial sorting scores, and further, the reward scores are determined for each search result entry according to the matching situation between the initial sorting and the expected sorting. It should be noted that the reward rules for determining the reward scores can also be set in combination with the online strategy.
[0044] In a possible implementation, the expected sorting is divided into a first sorting area and a second sorting area according to a preset rule, and the initial sorting is divided into the first sorting area and the second sorting area according to the preset rule; then, the step of using the scoring and evaluation model to match the initial sorting of the multiple search result entries with the expected sorting of the multiple search result entries to determine the reward scores for the respective initial sorting scores includes:
[0045] Determine a first reward score for the search result entries in the first sorted region of the initial sorting that are consistent with the first sorted region of the desired sorting;
[0046] Determine a second reward score for the search result entries in the first sorted region of the initial sorting that are inconsistent with the first sorted region of the desired sorting;
[0047] Determine a third reward score for the search result entries in the second sorted region of the initial sorting that are consistent with the second sorted region of the desired sorting;
[0048] Determine a fourth reward score for the search result entries in the second sorted region of the initial sorting that are inconsistent with the second sorted region of the desired sorting;
[0049] Wherein, the first reward score and the third reward score are positive numbers, and the second reward score and the fourth reward score are negative numbers.
[0050] Specifically, the sorting result can be divided into a first sorted region and a second sorted region according to preset rules such as the display number. For example, when N = 10, the region of Top1-5 can be used as the first sorted region, and Top6-10 can be used as the second sorted region. The search result entries within each sorted region can be the number that can be displayed to the user at one time. The first sorted region is the more concerned head region, and the accuracy of its sorting result greatly affects the user's search experience. Therefore, the initial sorting within the first sorted region can be evaluated first. Specifically, a first reward score can be determined for the search result entries in the first sorted region of the initial sorting that precisely match the first sorted region of the desired sorting, which can be a relatively high reward; a second reward score can be determined for the search result entries in the first sorted region of the initial sorting that do not match the first sorted region of the desired sorting, which can be a penalty score. Similarly, the initial sorting within the second sorted region is evaluated.
[0051] It should be noted that the reward levels set for the evaluation rules of the first sorted region and the second sorted region can be the same or different, that is, the first reward score and the third reward score can be the same or different, and the second reward score and the fourth reward score can be the same or different. Specifically, it can be set according to actual training requirements, and this application does not make any limitations on this.
[0052] It can be understood that, in addition to the above method of dividing and sorting the regions and evaluating them by region, it is also possible to directly determine the reward score for each initial sorting score according to the expected sorting from 1 to N based on the matching degree between the initial sorting and the expected sorting. In addition, N reward scores can be set according to the expected sorting from 1 to N, and specifically, they can be set according to actual training requirements. This application does not make any limitations in this regard.
[0053] S104: Determine the loss function of the multiple search result entries based on the initial sorting score and the reward score.
[0054] Since the reward score is determined based on the expected sorting, the reward score can represent the matching degree between the initial sorting score and the expected sorting, that is, it represents the accuracy of the initial sorting score output by the initial sorting model. Therefore, according to the initial sorting score and the reward score, the loss function is determined to train the initial scoring and sorting model using this loss function.
[0055] In a possible implementation manner, the initial sorting score can be normalized to obtain the initial sorting loss of the multiple search result entries. Further, according to the product of the initial sorting loss and the reward score, the loss function of the multiple search result entries is determined.
[0056] S105: Train the initial scoring and sorting model according to the loss function to obtain a target scoring and sorting model for re - sorting the multiple search result entries.
[0057] Using the loss function determined according to the initial sorting score and the reward score to train the initial scoring and sorting model enables the initial scoring and sorting model to be optimized in the direction of outputting the expected sorting, and finally obtain the target scoring and sorting model, improving the accuracy of the sorting result. This target scoring and sorting model is used to re - sort the multiple search result entries.
[0058] It can be understood that in the final online inference stage, the target scoring and sorting model can be used to achieve the re - sorting of the multiple search result entries for the target search term. Therefore, in a possible implementation manner, after S105, the following steps are further included:
[0059] S11: Use the target scoring and sorting model to score and sort the multiple search result entries of the target search term to obtain the target sorting of the target search term;
[0060] S12: Display the multiple search result entries according to the target sorting.
[0061] It can be seen that the target scoring and ranking model is used for online inference, while the scoring and evaluation model only takes effect in the offline training stage.
[0062] Specifically, after obtaining the target scoring and ranking model, use the target scoring and ranking model to score and rank multiple search result entries of the target search term, obtain the target ranking of the target search term, and display the multiple search result entries based on this target ranking.
[0063] In a possible implementation, after obtaining the target scoring and ranking model, the scoring and evaluation model can also be deleted, and only the target scoring and ranking model is retained in the subsequent online inference usage stage.
[0064] When performing online inference, due to the influence of time consumption, in a possible implementation, only the search result entries at the head can be rearranged, while the others can remain in the same order as in the sample training set.
[0065] It can be understood that the method of using the scoring and evaluation model to determine the reward score according to the expected ranking for the initial ranking score output by the initial scoring and ranking model realizes the purpose of unsupervised training of the initial scoring and ranking model. The training method of using the scoring and evaluation model and the initial scoring and ranking model finally obtains the target scoring and ranking model, which is a training method based on reinforcement learning. In a possible implementation, the initial scoring and ranking model and the scoring and evaluation model adopt the Actor-Critic model, where both Actor and Critic adopt the Transformer structure. The Actor-Critic model based on the Transformer structure of reinforcement learning, and the target scoring and ranking model determined by the above training method is a rearrangement model, which can effectively solve the problems of large mutual exclusion of policy fusion and lack of global perception, thereby greatly improving the user's search experience and the overall market distribution time, and improving the policy iteration efficiency on the rearrangement side.
[0066] It can be seen that the initial sorting scores corresponding to each are obtained by initially sorting and scoring multiple search result entries of the target search term through the initial scoring and sorting model. The scoring and evaluation model evaluates the initial sorting scores output by the initial scoring and sorting model according to the expected sorting of the multiple search result entries, and determines their respective reward scores. Further, a loss function of the multiple search result entries is determined according to the initial sorting scores and the reward scores, and the initial scoring and sorting model is trained for the sorting model according to the loss function, so as to obtain a target scoring and sorting model for re-sorting the multiple search result entries. It can be seen that the method of using the scoring and evaluation model to determine the reward scores according to the expected sorting for the initial sorting scores output by the initial scoring and sorting model realizes the purpose of unsupervised training of the initial scoring and sorting model, and the reward scores are determined based on the expected sorting. Therefore, the initial scoring and sorting model is trained for the sorting model by using the loss function determined according to the initial sorting scores and the reward scores, so that the initial scoring and sorting model can be optimized in the direction of outputting the expected sorting, and finally the target scoring and sorting model is obtained, improving the accuracy of the sorting result. The target scoring and sorting model is used to re-sort the multiple search result entries, so as to display the relevant search results of the target search term for the user according to the sorting result after re-sorting.
[0067] Figure 2 FIG. is a schematic framework diagram of a method for determining a search re-ranking model provided by an embodiment of the present application, which can execute the method for determining the search re-ranking model provided by the above embodiment, and specifically includes two parts as follows:
[0068] The first part is a training part including training the Actor and the Critic using a training sample set; the second part is online inference using the Actor. Among them, the Actor is a scoring and sorting model, and the initial scoring and sorting model is trained for the sorting model to obtain a target scoring and sorting model finally used for online inference; the Critic is a scoring and evaluation model, which is used to evaluate the initial scoring output by the initial scoring and sorting model, and determine the corresponding reward scores for the initial scoring output by the initial scoring and sorting model, so as to determine the loss function according to the initial scoring and the reward scores and train the initial scoring and sorting model for the sorting model.
[0069] Based on this, the method of using the Actor scoring and sorting model to complete the scoring and sorting, and using the Critic scoring and evaluation model to determine the reward scores according to the expected sorting for the initial sorting scores output by the Actor realizes the purpose of unsupervised training of the Actor. The trained Actor is used as the target scoring and sorting model, and in the subsequent online inference stage, the re-sorting of multiple retrieval result entries of the target retrieval term is realized, so as to display the relevant search results of the target search term for the user according to the sorting result after re-sorting.
[0070] Figure 3 This is the device structure diagram of a device for determining a search rearrangement model provided by an embodiment of the present application. The device includes an acquisition unit 301, a determination unit 302, and a training unit 303:
[0071] The acquisition unit 301 is configured to acquire a training sample set corresponding to a target search term; the training sample set includes the target search term and multiple search result entries of the target search term;
[0072] The determination unit 302 is configured to input the training sample set into an initial scoring and ranking model, and determine the initial ranking scores of the multiple search result entries respectively through the initial scoring and ranking model;
[0073] The determination unit 302 is further configured to input the training sample set including the target ranking label and the initial ranking scores into a scoring and evaluation model, and determine the reward scores of the initial ranking scores respectively through the scoring and evaluation model based on the target ranking label and the initial ranking scores; wherein, the target ranking label is used to identify the expected ranking of the multiple search result entries in the training sample set;
[0074] The determination unit 302 is further configured to determine the loss function of the multiple search result entries based on the initial ranking scores and the reward scores;
[0075] The training unit 303 is configured to perform ranking model training on the initial scoring and ranking model according to the loss function to obtain a target scoring and ranking model for performing search rearrangement on the multiple search result entries.
[0076] In a possible implementation manner, the determination unit is further configured to:
[0077] Input the training sample set into the initial scoring and ranking model, and determine the feature vectors of the multiple search result entries in the training sample set through the feature extraction layer of the initial scoring and ranking model;
[0078] Determine the initial ranking scores of the multiple search result entries respectively according to the feature vectors.
[0079] In a possible implementation manner, the determination unit is further configured to:
[0080] Match the initial ranking of the multiple search result entries with the expected ranking of the multiple search result entries through the scoring and evaluation model, and determine the reward scores of the initial ranking scores respectively; the initial ranking of the multiple search result entries is determined based on the initial ranking scores.
[0081] In a possible implementation, the determining unit is further configured to:
[0082] Divide the expected sorting into a first sorting area and a second sorting area according to a preset rule, and divide the initial sorting into the first sorting area and the second sorting area according to the preset rule;
[0083] Determine a first reward score for the search result entries that are consistent in the first sorting area of the initial sorting and the first sorting area of the expected sorting;
[0084] Determine a second reward score for the search result entries that are inconsistent in the first sorting area of the initial sorting and the first sorting area of the expected sorting;
[0085] Determine a third reward score for the search result entries that are consistent in the second sorting area of the initial sorting and the second sorting area of the expected sorting;
[0086] Determine a fourth reward score for the search result entries that are inconsistent in the second sorting area of the initial sorting and the second sorting area of the expected sorting;
[0087] Wherein, the first reward score and the third reward score are positive numbers, and the second reward score and the fourth reward score are negative numbers.
[0088] In a possible implementation, the determining unit is further configured to:
[0089] Normalize the initial sorting scores to obtain the initial sorting loss of the multiple search result entries;
[0090] Determine the loss function of the multiple search result entries according to the initial sorting loss multiplied by the reward score.
[0091] In a possible implementation, the training unit is further configured to:
[0092] Use the time difference method to train the scoring and evaluation model.
[0093] In a possible implementation, the determining unit is further configured to:
[0094] Use the target scoring and sorting model to score and sort the multiple search result entries of the target search term to obtain the target sorting of the target search term;
[0095] Display the multiple search result entries according to the target sorting.
[0096] In a possible implementation manner, the initial scoring and ranking model and the scoring evaluation model adopt the Actor-Critic model.
[0097] It can be seen that through the initial scoring and ranking model, initial ranking scores corresponding to each are obtained by initially ranking and scoring multiple search result entries of the target search term. Through the scoring evaluation model, the initial ranking scores output by the initial scoring and ranking model are evaluated according to the expected ranking of the multiple search result entries, and respective reward scores are determined. Further, a loss function of the multiple search result entries is determined according to the initial ranking scores and the reward scores, and the initial scoring and ranking model is trained for the ranking model according to this loss function to obtain a target scoring and ranking model for search re-ranking of the multiple search result entries. It can be seen that by using the scoring evaluation model to determine the reward scores based on the expected ranking for the initial ranking scores output by the initial scoring and ranking model, the purpose of unsupervised training of the initial scoring and ranking model is achieved, and the reward scores are determined based on the expected ranking. Therefore, by using the loss function determined according to the initial ranking scores and the reward scores to train the initial scoring and ranking model for the ranking model, the initial scoring and ranking model can be optimized in the direction of outputting the expected ranking, and finally the target scoring and ranking model is obtained, improving the accuracy of the ranking result. The target scoring and ranking model is used to perform search re-ranking on the multiple search result entries so as to display relevant search results of the target search term for the user according to the ranking result after search re-ranking.
[0098] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0099] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0100] The above has introduced in detail a method and device for determining a search rearrangement model provided by an embodiment of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application. At the same time, for those of ordinary skill in the art, according to the method of the present application, there will be changes in the specific implementation manner and application scope.
[0101] In summary, the content of this specification should not be construed as a limitation to the present application. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Moreover, based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
Claims
1. A method for determining a search rearrangement model, characterized in that The method includes: Obtaining a training sample set corresponding to a target search term; the training sample set includes the target search term and multiple search result entries of the target search term; Inputting the training sample set into an initial scoring and ranking model, and determining respective initial ranking scores of the multiple search result entries through the initial scoring and ranking model; Inputting the training sample set including target ranking labels and the initial ranking scores into a scoring evaluation model, and determining respective reward scores of the initial ranking scores through the scoring evaluation model based on the target ranking labels and the initial ranking scores; wherein, the target ranking labels are used to identify the expected rankings of the multiple search result entries in the training sample set; Determining a loss function of the multiple search result entries based on the initial ranking scores and the reward scores; Training the initial scoring and ranking model according to the loss function to obtain a target scoring and ranking model for re-ranking the multiple search result entries; Dividing the expected rankings into a first ranking region and a second ranking region according to a preset rule, and dividing the initial rankings of the multiple search result entries into the first ranking region and the second ranking region according to the preset rule; the initial rankings of the multiple search result entries are determined based on the initial ranking scores; The determining, through the scoring evaluation model, respective reward scores of the initial ranking scores based on the target ranking labels and the initial ranking scores includes: Determining a first reward score for search result entries in the first ranking region of the initial ranking that are consistent with those in the first ranking region of the expected ranking; Determining a second reward score for search result entries in the first ranking region of the initial ranking that are inconsistent with those in the first ranking region of the expected ranking; Determining a third reward score for search result entries in the second ranking region of the initial ranking that are consistent with those in the second ranking region of the expected ranking; Determining a fourth reward score for search result entries in the second ranking region of the initial ranking that are inconsistent with those in the second ranking region of the expected ranking; Wherein, the first reward score and the third reward score are positive numbers, and the second reward score and the fourth reward score are negative numbers.
2. The method according to claim 1, wherein The inputting the training sample set into the initial scoring and ranking model, and determining respective initial ranking scores of the multiple search result entries through the initial scoring and ranking model includes: Inputting the training sample set into the initial scoring and ranking model, and determining feature vectors of the multiple search result entries in the training sample set through a feature extraction layer of the initial scoring and ranking model; Determining respective initial ranking scores of the multiple search result entries according to the feature vectors.
3. The method according to claim 1, wherein It further includes: Performing normalization processing on the initial ranking scores to obtain initial ranking losses of the multiple search result entries; Then, the determining the loss function of the multiple search result entries based on the initial ranking scores and the reward scores includes: Determine a loss function for the multiple search result entries by multiplying the initial sorting loss by the reward score.
4. The method according to any one of claims 1 to 3, characterized in that It further includes: Use the temporal difference method to train the scoring and evaluation model.
5. The method according to any one of claims 1 to 3, characterized in that, After obtaining the target scoring and sorting model for re-sorting the multiple search result entries, it further includes: Use the target scoring and sorting model to score and sort the multiple search result entries of the target search term to obtain the target sorting of the target search term; Display the multiple search result entries according to the target sorting.
6. The method according to any one of claims 1-3, characterized in that, The initial scoring and sorting model and the scoring and evaluation model adopt the Actor-Critic model.
7. An apparatus for determining a search rearrangement model, characterized in that The device includes an acquisition unit, a determination unit, and a training unit: The acquisition unit is used to acquire a training sample set corresponding to a target search term; the training sample set includes the target search term and multiple search result entries of the target search term; The determination unit is used to input the training sample set into the initial scoring and sorting model, and determine the initial sorting scores of the multiple search result entries respectively through the initial scoring and sorting model; The determination unit is further used to input the training sample set including the target sorting label and the initial sorting score into the scoring and evaluation model, and determine the reward scores of the initial sorting scores respectively based on the target sorting label and the initial sorting score through the scoring and evaluation model; wherein, the target sorting label is used to identify the expected sorting of the multiple search result entries in the training sample set; The determination unit is further used to determine a loss function for the multiple search result entries based on the initial sorting score and the reward score; The training unit is used to train the initial scoring and sorting model according to the loss function to obtain a target scoring and sorting model for re-sorting the multiple search result entries; The determination unit is further used to: Divide the expected sorting into a first sorting area and a second sorting area according to a preset rule, and divide the initial sorting of the multiple search result entries into the first sorting area and the second sorting area according to the preset rule; the initial sorting of the multiple search result entries is determined based on the initial sorting score; Determine a first reward score for the search result entries that are consistent in the first sorting area of the initial sorting and the first sorting area of the expected sorting; Determine a second reward score for the search result entries that are inconsistent in the first sorting area of the initial sorting and the first sorting area of the expected sorting; Determine a third reward score for the search result entries that are consistent in the second sorting area of the initial sorting and the second sorting area of the expected sorting; Determine a fourth reward score for the search result entries that are inconsistent in the second sorting area of the initial sorting and the second sorting area of the expected sorting; Wherein, the first reward score and the third reward score are positive numbers, and the second reward score and the fourth reward score are negative numbers.
Citation Information
Patent Citations
Method for generating ranking model, method for ranking search results, apparatus and apparatus
CN109299344A
Object pushing method and device, computer equipment and storage medium
CN110413893A
Method and device for displaying target object sequence to target user
CN111915414A