Download rate estimation model training method and device, download rate estimation model searching method and device, equipment and medium

By training the download rate estimate model and using reinforcement learning mechanism to process user search behavior data, the problem of inaccurate sorting of search results in the existing technology is solved, and the accuracy and user experience of search results are improved.

CN119961506APending Publication Date: 2025-05-09BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311484717.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The search results sorting method of existing app stores is based on preset standards and may not accurately reflect users' search needs and affect user experience.

Method used

By obtaining the search behavior data of sample users under the search path, the reinforcement learning mechanism based on human feedback trains the download rate estimate model to sort search results.

Benefits of technology

It improves the estimate accuracy of the download rate estimate model, makes the search results ranked first more in line with users' search needs, and improves the accuracy and user experience of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961506A_ABST
    Figure CN119961506A_ABST
Patent Text Reader

Abstract

The invention relates to a download rate prediction model training method and device, a search method and device, equipment and a medium. The method comprises the steps of obtaining search behavior data of a sample user under a search path, wherein the search path is used for representing a switching path between search scenes; and according to the search behavior data, performing model training based on a reinforcement learning mechanism fed back by human beings to obtain a download rate estimation model. Therefore, modeling can be carried out on the whole user search behavior link, decoupling can be carried out aiming at a search service scene, systematic modeling is achieved, then the estimation accuracy of the download rate estimation model is improved, the search result ranked in the front fits the search requirement of the user, the accuracy of the output search result and the user experience are improved, and the user experience is improved. And the downloading rate of the search engine is further improved. Besides, model training is carried out based on a reinforcement learning mechanism of human feedback, iterative optimization can be carried out on the download rate estimation model, and the estimation accuracy of the download rate estimation model is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a download rate prediction model training method, search method, device, equipment and medium. Background Art

[0002] As terminals are increasingly widely used, more and more applications are installed on terminals including mobile phones and tablets. Users generally search for the applications they need to use through the search interface of the application store. In related technologies, the application store will query a series of results related to the search content based on the search content entered by the user, and estimate the click-through rate (CTR) of these results, and then evaluate the estimated results according to the preset standards, and then sort these results, and output the sorted results to the user for viewing. However, the search results output after sorting the information based on the preset standards have limitations. The search results ranked in the front may not fit the user's actual search needs, affecting the user's search experience. Summary of the invention

[0003] In order to overcome the problems existing in the related art, the present disclosure provides a download rate prediction model training method, search method, device, equipment and medium.

[0004] According to a first aspect of an embodiment of the present disclosure, a download rate prediction model training method is provided, comprising: obtaining search behavior data of sample users under a search path, wherein the search path is used to characterize a switching path between search scenarios; and performing model training based on a reinforcement learning mechanism of human feedback according to the search behavior data to obtain a download rate prediction model.

[0005] Optionally, the search behavior data includes: the first search sub-behavior data of the sample user from the search results page to the search details page, wherein the first search sub-behavior data includes the historical search results of the search results page and the actual download status of the sample user on the search results page; the second search sub-behavior data of the sample user from the search suggestion page to the search results page and the search details page, wherein the second search sub-behavior data includes the suggested results of the search suggestion page and the actual download status of the sample user on the search suggestion page; the download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after triggering the search.

[0006] Optionally, according to the search behavior data, a reinforcement learning mechanism based on human feedback is used to perform model training to obtain a download rate prediction model, including: according to the first search sub-behavior data, a basic model is supervised fine-tuned to obtain a supervised fine-tuned model; according to the second search sub-behavior data and the download behavior sequence, the supervised fine-tuned model is fine-tuned by a reinforcement learning method to obtain a download rate prediction model.

[0007] Optionally, the basic model is a location deviation-aware learning model including a multi-gated hybrid expert network and a location network, and the historical search results and the recommended results are both composed of reference applications; the supervised fine-tuning model is fine-tuned by a reinforcement learning method based on the second search sub-behavior data and the download behavior sequence to obtain a download rate prediction model, including: obtaining the first feature information of the sample user, the search keyword corresponding to the search behavior data, and the second feature information of the search keyword; for each of the reference applications in the recommended results, obtaining the third feature information of the reference application; generating a feature vector of the reference application based on the first feature information, the search keyword, the second feature information, the reference application, the third feature information and the download behavior sequence; determining the download label of the sample user for the reference application based on the actual download situation of the sample user on the search suggestion page; and fine-tuning the multi-gated hybrid expert network of the supervised fine-tuning model by a reinforcement learning method based on each of the feature vectors and each of the download labels to obtain a download rate prediction model.

[0008] Optionally, the multi-gated hybrid expert network includes a first gating network, a second gating network, multiple expert networks, a first tower network, a second tower network, a first processing unit, a second processing unit and a third processing unit; the multi-gated hybrid expert network of the supervised fine-tuning model is fine-tuned by a reinforcement learning method according to each of the feature vectors and each of the download labels to obtain a download rate prediction model, including: taking the feature vector of the reference application as the input of the first gating network, the second gating network, and each of the expert networks, taking the output of each of the expert networks and the output of the first gating network as the input of the first processing unit, and taking the output of the first processing unit and the first sub-feature vector as the first tower network. input, using the output of each of the expert networks and the output of the second gating network as the input of the second processing unit, using the output of the second processing unit and the second sub-feature vector as the input of the second tower network, using the output of the first tower network and the output of the second tower network as the input of the third processing unit, and using the download label of the sample user for the reference application as the target output of the third processing unit, performing reinforcement learning training on the multi-gated hybrid expert network of the supervised fine-tuning model to obtain a download rate estimation model, wherein the first sub-feature vector is a feature vector generated according to the reference application and the third feature information, and the second sub-feature vector is a feature vector generated according to the download behavior sequence.

[0009] Optionally, the search behavior data includes search behavior data of the sample user under two search engines.

[0010] According to the second aspect of an embodiment of the present disclosure, a search method is provided, comprising: in response to receiving a search request, obtaining target search results matching the search request, wherein the target search results include multiple target applications; determining the estimated download rate of each of the target applications respectively through a pre-trained download rate prediction model, wherein the download rate prediction model is trained based on the download rate prediction model training method provided by the first aspect of the present disclosure; sorting the target search results according to the estimated download rate of each of the target applications, and displaying the sorted target search results.

[0011] According to the third aspect of an embodiment of the present disclosure, a download rate prediction model training device is provided, including: a first acquisition module, configured to obtain search behavior data of sample users under a search path, wherein the search path is used to characterize a switching path between search scenarios; a training module, configured to perform model training based on a reinforcement learning mechanism of human feedback according to the search behavior data to obtain a download rate prediction model.

[0012] According to the fourth aspect of an embodiment of the present disclosure, a search device is provided, comprising: a second acquisition module, configured to, in response to receiving a search request, acquire target search results matching the search request, wherein the target search results include multiple target applications; an estimation module, configured to determine the estimated download rate of each of the target applications respectively through a pre-trained download rate estimation model, wherein the download rate estimation model is trained based on the download rate estimation model training method provided in the first aspect of the present disclosure; and a display module, configured to sort the target search results according to the estimated download rate of each of the target applications, and display the sorted target search results.

[0013] According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to: implement the steps of the download rate prediction model training method provided in the first aspect of the present disclosure or the steps of the search method provided in the second aspect of the present disclosure when executing the executable instructions.

[0014] According to the sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the download rate prediction model training method provided in the first aspect of the present disclosure or the steps of the search method provided in the second aspect of the present disclosure are implemented.

[0015] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, the search behavior data of sample users under the search path is obtained, wherein the search path is used to characterize the switching path between search scenarios; then, based on the search behavior data, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model. In this way, the entire user search behavior link can be modeled, so that the search business scenario can be decoupled to achieve systematic modeling, thereby improving the estimation accuracy of the download rate estimation model, so that the top search results are consistent with the user's search needs, improving the accuracy of the output search results and user experience, and thereby improving the download rate of the search engine. In addition, by training the model based on a reinforcement learning mechanism of human feedback, the download rate estimation model can be iteratively optimized, further improving the estimation accuracy of the download rate estimation model.

[0016] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0018] Figure 1 It is a schematic diagram of a search suggestion page according to an exemplary embodiment.

[0019] Figure 2 The figure is a schematic diagram of a search result page according to an exemplary embodiment.

[0020] Figure 3 The figure is a schematic diagram of a search details page according to an exemplary embodiment.

[0021] Figure 4 It is a flowchart of a download rate estimation model training method according to an exemplary embodiment.

[0022] Figure 5 It is a structural diagram of a basic model according to an exemplary embodiment.

[0023] Figure 6 is a flow chart showing a search method according to an exemplary embodiment.

[0024] Figure 7 It is a block diagram of a download rate prediction model training device according to an exemplary embodiment.

[0025] Figure 8 The figure is a block diagram of a search device according to an exemplary embodiment.

[0026] Fig. 9 It is a block diagram of a device for training a download rate prediction model according to an exemplary embodiment.

[0027] Fig.10 It is a block diagram of a device for searching according to an exemplary embodiment. DETAILED DESCRIPTION

[0028] Before describing the embodiments of the present disclosure, the application scenarios of the present application are first described. The search scenarios of the present disclosure can be divided into: search suggestion (Suggestion, abbreviated as Sug) page, search results page and search details page. Among them, the Sug page refers to the search suggestion page of the search engine, which is used to display search suggestions (i.e., suggested results) and prompts; the search results page refers to the search results page of the search engine, which is used to display search results; the search details page refers to the search details page of the search engine, which is used to display detailed information of a search item in the search results or suggestion results. Among them, the suggestion results and search results may include one or more search items.

[0029] Specifically, when a user has a search need, he can enter the search keyword in the search box on the search homepage, and the page will jump to the Sug page. If the search item that the user wants to download (for example, an application) is included in the recommended results on the Sug page, the user can directly click the "Install" or "Download" button to download and install it. If the user clicks the search button on the Sug page, the page will jump to the search results page. Of course, the user can also jump to the search details page of the corresponding search item by clicking any search item in the recommended results on the Sug page (instead of the "Install" or "Download" button).

[0030] In addition, if the page jumps from the Sug page to the search results page, when the search item that the user wants to download exists on the search results page, the user can download and install it by clicking the "Install" or "Download" button. Of course, the user can also jump to the search details page of the corresponding search item by clicking any search item in the search results of the search results page (instead of the "Install" or "Download" button).

[0031] If the page jumps from the Sug page or the search results page to the search details page of a certain search item, when the user wants to download the search item, the user can download and install it by clicking the "Install" or "Download" button on the search details page.

[0032] If the user does not find the desired search item in the above search process, the above search process will be repeated to search again.

[0033] For example, taking the search engine as the game center, the user enters the search keyword "King of Glory" in the search box on the search homepage, and the page will jump to Figure 1 The Sug page shown in FIG. 1 includes four games (i.e., search items). After that, the user clicks the search button on the Sug page and the page jumps to Figure 2 Next, the user clicks the search item "Honor of Kings" in the search results of the search results page (rather than the "Install" or "Download" button) to jump to the Figure 3 The search details page of "Honor of Kings" is shown. At this time, if the user wants to download the game "Honor of Kings", he can click the "Install" button to download and install it.

[0034] The exemplary embodiments will be described in detail below, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0035] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.

[0036] Figure 4 FIG. 1 is a flowchart of a download rate estimation model training method according to an exemplary embodiment. Figure 4 As shown, the download rate estimation model training method may include the following S101 and S102.

[0037] In S101, the search behavior data of the sample users in the search path is obtained.

[0038] In the present disclosure, the search path is used to characterize the switching path between search scenarios. Among them, the search scenarios can be divided into: Sug page, search result page and search details page. Accordingly, the search path can include: Sug page→search result page, Sug page→search details page, search result page→search details page.

[0039] The above-mentioned search behavior data may include: the first search sub-behavior data of the sample user from the search results page to the search details page, wherein the first search sub-behavior data includes the historical search results of the search results page and the actual download status of the sample user on the search results page; the second search sub-behavior data of the sample user from the search suggestion page to the search results page and the search details page, wherein the second search sub-behavior data includes the suggested results of the search suggestion page and the actual download status of the sample user on the search suggestion page; the download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after triggering the search.

[0040] Among them, the actual download situation of the sample user on the search results page is used to characterize whether the sample user has download behavior on the search results page, for example, the sample user has downloaded a search item on the search results page, or the sample user has not downloaded on the search results page, but has jumped from the search results page to the search details page. The actual download situation of the sample user on the search suggestion page is used to characterize whether the sample user has download behavior on the Sug page, for example, the sample user has downloaded a search item on the Sug page, or the sample user has not downloaded on the Sug page, but has jumped from the Sug page to the search details page or the search results page.

[0041] The download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after the search is triggered refers to the sequence of search items downloaded by the sample user on the search results page and the search suggestion page within a preset time after the user opens the search engine.

[0042] For example, within 5 minutes of opening the game center, the user first downloads Game A and Game E in sequence on the Sug page by searching for keyword C1; then, he downloads Game D on the search results page by searching for keyword C2. The user's download behavior sequence within 5 minutes after triggering the search is: Game A, Game E, Game D.

[0043] In addition, when a sample user searches and downloads through a search engine, the search behavior data of the sample user in the corresponding search path can be recorded as a training sample.

[0044] In S102, according to the search behavior data of sample users in the search path, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model.

[0045] In the present disclosure, the search service is an active behavior, and therefore, the entire user search behavior link can be modeled using a reinforcement learning mechanism based on human feedback (RLHF).

[0046] The search behavior data may include the search behavior data of sample users under two search engines, that is, the training samples come from data in two different fields. In this way, the download rate prediction model can learn the common characteristics of data in two different fields, improve the applicability of the model and the accuracy of download rate prediction. In addition, a larger amount of training data can further improve the accuracy of the model's download rate prediction.

[0047] For example, the search behavior data may include the search behavior data of the sample users in an app store and the search behavior data of the sample users in a game center.

[0048] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, the search behavior data of sample users under the search path is obtained, wherein the search path is used to characterize the switching path between search scenarios; then, based on the search behavior data, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model. In this way, the entire user search behavior link can be modeled, so that the search business scenario can be decoupled to achieve systematic modeling, thereby improving the estimation accuracy of the download rate estimation model, so that the top search results are consistent with the user's search needs, improving the accuracy of the output search results and user experience, and thereby improving the download rate of the search engine. In addition, by training the model based on a reinforcement learning mechanism of human feedback, the download rate estimation model can be iteratively optimized, further improving the estimation accuracy of the download rate estimation model.

[0049] The following is a detailed description of the specific implementation method of performing model training based on search behavior data and human feedback-based reinforcement learning mechanism in S102 to obtain a download rate estimation model. Specifically, it can be achieved through the following steps (1) and (2):

[0050] Step (1): Based on the first search sub-behavior data, the basic model is fine-tuned in a supervised manner to obtain a supervised fine-tuning model.

[0051] Step (2): Based on the second search sub-behavior data and the download behavior sequence, the supervised fine-tuning model is fine-tuned by a reinforcement learning method to obtain a download rate estimation model.

[0052] In the present disclosure, a reinforcement learning mechanism based on human feedback can be used to train the download rate prediction model in a similar manner to the InstrcutGPT model training. Specifically, first, the first search sub-behavior data of the sample user from the search results page to the search details page can be used to perform supervised fine-tuning training on the basic model to obtain a supervised fine-tuning (SFT) model; then, the second search sub-behavior data of the sample user from the search suggestion page to the search results page and the search details page, as well as the download behavior sequence of the sample user on the search results page and the search suggestion page, are used to fine-tune and strengthen the SFT model to obtain a download rate prediction model. Here, the original reward model (Reward Model) and reinforcement learning (RL) training in InstrcutGPT are combined.

[0053] Among them, Figure 5 As shown, the above-mentioned basic model can be a Position-bias Aware Learning framework (PAL) model including a multi-gate mixture of experts (Multi-gate Mixture-of-Experts, MMoE) network and a position (Position) network. The use of the PAL model can avoid the exposure bias problem caused by the search item position.

[0054] The specific implementation method of performing supervised fine-tuning training on the basic model according to the first search sub-behavior data in the above step (1) is described in detail below. Specifically, it can be achieved by the following steps (11) to (16):

[0055] Step (11): Obtain the first characteristic information of the sample user, the search keyword corresponding to the search behavior data, and the second characteristic information of the search keyword.

[0056] In the present disclosure, the historical search results of the search results page and the suggested results of the Sug page are both composed of reference applications, that is, the search results page and the Sug page may contain one or more reference applications.

[0057] The first characteristic information of the sample user is the basic attribute information of the user, which may include the city where the user is located, the model of the user's terminal, etc. The second characteristic information of the search keyword may include the pinyin corresponding to the search keyword, the historical download volume of the search result corresponding to the search keyword, etc.

[0058] Step (12): For each reference application in the historical search results, obtain the third feature information of the reference application.

[0059] In the present disclosure, the third characteristic information of the reference application may include information such as a brief introduction and tag words of the reference application.

[0060] Step (13): Generate a feature vector of the reference application in the search results based on the first feature information, the search keyword, the second feature information, the reference application and the third feature information.

[0061] In the present disclosure, the first feature information, search keywords, the second feature information, the reference application and the third feature information can be vectorized separately to obtain multiple first vectors. Then, the multiple first vectors are concatenated in sequence to obtain the feature vector of the reference application in the search results.

[0062] Step (14): According to the actual downloading situation of the sample users on the search results page, determine the download tags of the sample users for the reference application.

[0063] In the present disclosure, the search results page may include one or more reference applications. If the actual download situation of the sample user on the search results page indicates that the sample user has downloaded the reference application on the search results page, then the download label of the sample user for the reference application is determined to be 1 (that is, the download rate is 1); if the actual download situation of the sample user on the search results page indicates that the sample user has not downloaded the reference application on the search results page (that is, the sample user may have downloaded other reference applications on the search results page, or may have jumped from the search results page to the search details page), then the download label of the sample user for the reference application is determined to be 0 (that is, the download rate is 0).

[0064] Step (15): Obtain the location information of the reference application in the search results page.

[0065] For example, the location information is used to characterize the ranking of the reference application in the historical search results of the search result page.

[0066] Step (16): Based on the feature vector of each reference application in the search results, the download tags of the sample users for each reference application in the historical search results, and the location information of each reference application in the search results page, the basic model is fine-tuned in a supervised manner to obtain the SFT model.

[0067] In the present disclosure, for each reference application in the historical search results, the feature vector of the reference application, the download tag of the reference application, and the location information of the reference application in the search result page can be used as a training sample to perform supervised fine-tuning training on the basic model. For example, if the historical search results include five reference applications, then the user can obtain five training samples at one time for the current sample.

[0068] Specifically, if Figure 5 As shown, for each training sample in the historical search results, the location information of the reference application in the training sample in the search results page can be input into the location network to obtain the probability Proseen that the reference application is seen by the sample user on the search results page; at the same time, the feature vector of the reference application in the search results is input into the MMoE network to obtain the probability pCTR that the reference application in the search results is downloaded after being seen by the sample user; then, the product of Proseen and pCTR is determined as the estimated download rate of the reference application in the historical search results by the sample user; then, according to the difference between the estimated download rate of the reference application in the historical search results by the sample user and the corresponding download label in the training sample, the model parameters of the basic model are updated; then, new training data is re-acquired and the model training is continued until the first training cutoff condition is reached.

[0069] In one embodiment, the first training cutoff condition may be that the number of training times reaches a first preset number of times, and the first preset number of times can be set according to the actual usage scenario. When the number of training times reaches the first preset number of times, it can be determined that the number of training times is sufficient and the basic model can learn enough effective features.

[0070] In another embodiment, the first training cutoff condition may be that the target loss of the base model is less than a first preset threshold, and the first preset threshold may be set according to the actual usage scenario. When the target loss of the base model is less than the first preset threshold, it can be considered that the estimation accuracy of the base model meets the requirements and can accurately estimate the download rate.

[0071] The following describes in detail the specific implementation method of fine-tuning the supervised fine-tuning model based on the second search sub-behavior data and the download behavior sequence in the above step (2) to obtain the download rate estimation model through the reinforcement learning method.

[0072] Specifically, it can be achieved by following the steps (21) to (25):

[0073] Step (21): Obtain the first characteristic information of the sample user, the search keyword corresponding to the search behavior data, and the second characteristic information of the search keyword.

[0074] Step (22): For each reference application in the recommendation result, obtain the third feature information of the reference application.

[0075] Step (23): Generate a feature vector of the reference application based on the first feature information, the search keyword, the second feature information, the reference application, the third feature information, and the download behavior sequence.

[0076] In the present disclosure, the first feature information, search keywords, the second feature information, the reference application, the third feature information and the download behavior sequence can be vectorized separately to obtain multiple second vectors. Thereafter, the multiple second vectors can be concatenated in sequence to obtain the feature vector of the reference application in the recommended results.

[0077] Step (24): Based on the actual downloading situation of the sample users on the search suggestion page, determine the download tags of the sample users for the reference application.

[0078] In the present disclosure, a Sug page may include one or more reference applications. If the actual download situation of the sample user on the Sug page indicates that the sample user has downloaded the reference application on the Sug page, then the download label of the sample user for the reference application is determined to be 1 (i.e., the download rate is 1); if the actual download situation of the sample user on the Sug page indicates that the sample user has not downloaded the reference application on the Sug page (i.e., the sample user may have downloaded other reference applications on the Sug page, or may have jumped from the Sug page to the search details page or the search results page), then the download label of the sample user for the reference application is determined to be 0 (i.e., the download rate is 0).

[0079] Step (25): According to each feature vector and each download label, the multi-gated hybrid expert network of the supervised fine-tuning model is fine-tuned by a reinforcement learning method to obtain a download rate estimation model.

[0080] In the present disclosure, for each reference application in the recommended results of the Sug page, the feature vector of the reference application and the download label of the reference application can be used as a training sample to fine-tune the multi-gated hybrid expert network of the SFT model through a reinforcement learning method, thereby obtaining a download rate prediction model. Here, only the model parameters of the multi-gated hybrid expert network of the SFT model are fine-tuned, and the model parameters of the location network of the SFT model remain unchanged.

[0081] In addition, the multi-gated hybrid expert network of the SFT model can be fine-tuned through reinforcement learning methods such as the Proximal Policy Optimization (PPO) algorithm, Trust Region Policy Optimization (TRPO), Markov Decision Process (MDP), and Deep Deterministic Policy Gradient (DDPG).

[0082] like Figure 5 As shown, the multi-gated hybrid expert network may include a first gating network, a second gating network, a plurality of expert networks ( Figure 5 Three expert networks are used for example, specifically including a first expert network, a second expert network and a third expert network), a first tower network, a second tower network, a first processing unit, a second processing unit and a third processing unit, wherein one end of the first processing unit is connected to the first gating network and each expert network respectively, the other end of the first processing unit is connected to the first tower network, one end of the second processing unit is connected to the second gating network and each expert network respectively, the other end of the second processing unit is connected to the second tower network, and the third processing unit is connected to the location network, the first tower network and the second tower network respectively.

[0083] Among them, the structures of multiple expert networks are the same, the structures of the first gated network and the second gated network are the same. For example, the first gated network, the second gated network, each expert network, the first tower network, and the second tower network are all deep neural networks (Deep Nueral Network, DNN).

[0084] In combination with the specific structure of the multi-gated hybrid expert network, the specific implementation method of fine-tuning the multi-gated hybrid expert network of the supervised fine-tuning model according to each feature vector and each download label in the above step (25) by using the reinforcement learning method to obtain the download rate estimation model is described in detail.

[0085] In one embodiment, the multi-gated hybrid expert network of the supervised fine-tuning model can be trained by reinforcement learning by using the feature vector of the reference application in the recommendation result as the input of the first gating network, the second gating network, and each expert network, using the output of each expert network and the output of the first gating network as the input of the first processing unit, using the output of the first processing unit and the first sub-feature vector as the input of the first tower network, using the output of each expert network and the output of the second gating network as the input of the second processing unit, using the output of the second processing unit and the second sub-feature vector as the input of the second tower network, using the output of the first tower network and the output of the second tower network as the input of the third processing unit, and using the download label of the sample user for the reference application as the target output of the third processing unit to obtain a download rate prediction model.

[0086] The first sub-feature vector is a feature vector generated according to the reference application and the third feature information, that is, the first sub-feature vector is a concatenated vector of the second vector corresponding to the reference application and the second vector corresponding to the third feature information. The second sub-feature vector is a feature vector generated according to the download behavior sequence, that is, the second feature vector is the second vector corresponding to the download behavior sequence.

[0087] Specifically, for each training sample corresponding to the recommendation result: each expert network is used to further extract features according to the feature vector of the reference application in the training sample; the first gating network and the second gating network are used to assign weights to each expert network according to the feature vector of the reference application; the first processing unit is used to calculate the weighted sum of the features of the output of each expert network according to the weight assigned to each expert network by the first gating network, so as to obtain a first fused feature; the second processing unit is used to calculate the weighted sum of the features of the output of each expert network according to the weight assigned to each expert network by the second gating network, so as to obtain a second fused feature; the first tower network is used to perform deeper feature extraction on the first fused feature and the first sub-feature vector, so as to obtain a reference application in the recommendation result that is downloaded after being seen by the sample user. The first probability pCTR1, the second tower network is used to perform deeper feature extraction on the second fusion feature and the second sub-feature vector to obtain the second probability pCTR2 in the recommendation result that the reference application is downloaded after being seen by the sample user; the third processing unit is used to determine the estimated download rate of the sample user for the reference application in the recommendation result according to pCTR1 and pCTR2. For example, the product of pCTR1 and pCTR2 can be determined as the estimated download rate of the sample user for the reference application in the recommendation result; thereafter, the model parameters of the MMoE network of the basic model can be updated according to the difference between the estimated download rate of the sample user for the reference application in the recommendation result and the corresponding download label, and specifically the model parameters of the first gated network, the second gated network, each expert network, the first tower network, and the second tower network are updated.

[0088] Afterwards, new training data is acquired again and the MMoE network training is continued until the second training cutoff condition is reached.

[0089] In one embodiment, the second training cutoff condition may be that the number of trainings reaches a second preset number of times, and the second preset number of times may be set according to the actual usage scenario. When the number of trainings reaches the second preset number of times, it can be determined that the number of trainings is sufficient and the MMoE network can learn sufficient effective features.

[0090] In another embodiment, the second training cutoff condition may be that the target loss of the MMoE network is less than a second preset threshold, and the second preset threshold may be set according to an actual usage scenario. When the target loss of the MMoE network is less than the second preset threshold, it can be considered that the estimation accuracy of the MMoE network meets the requirements and the download rate can be accurately estimated.

[0091] In the above embodiment, the first tower network can be used as a reward model for the download rate prediction model, wherein the first sub-feature vector is determined based on the second search sub-behavior data from the search suggestion page to the search results page and the search details page, and the second sub-feature vector is determined based on the download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after the search is triggered. In this way, the first tower network can use the feature information of the second search sub-behavior data from the search suggestion page to the search results page and the search details page (i.e., the first sub-feature vector) to generalize the general word search of the SFT model, and the second tower network can use the feature information of the download behavior sequence (i.e., the second sub-feature vector) to accurately strengthen the SFT model, thereby improving the model effect of the download rate prediction model.

[0092] In combination with the specific structure of the multi-gated hybrid expert network, the specific implementation method of the supervised fine-tuning training of the basic model in the above step (16) is described in detail according to the feature vector of each reference application in the search results, the download tag of the sample user for each reference application in the historical search results, and the location information of each reference application in the search results page.

[0093] Specifically, the first tower network and the second tower network correspond to the above two search engines respectively. Figure 5 As shown, for each training sample corresponding to the historical search results: a location network is used to predict the probability Proseen1 that the reference application is seen by the sample user on the search results page based on the location information of the reference application in the search results page in the training sample; each expert network is used to further extract features based on the feature vector of the reference application in the search results.

[0094] When the training sample comes from the search engine corresponding to the first tower network, the first gating network is used to assign weights to each expert network according to the feature vector of the reference application; the first processing unit is used to calculate the weighted sum of the features of the output of each expert network according to the weights assigned to each expert network by the first gating network to obtain the third fusion feature; the first tower network is used to perform a deeper feature extraction on the third fusion feature to obtain the third probability pCTR3 that the reference application in the historical search results is downloaded after being seen by the sample user; the third processing unit is used to determine the estimated download rate of the sample user for the reference application in the historical search results according to Proseen1 and pCTR3. For example, the product of Proseen1 and pCTR3 can be determined as the estimated download rate of the sample user for the reference application in the historical search results. At this time, the second gating network, the second processing unit and the second tower network do not perform data processing.

[0095] When the training sample comes from the search engine corresponding to the second tower network, the second gating network is used to assign weights to each expert network according to the feature vector of the reference application; the second processing unit is used to calculate the weighted sum of the features of the output of each expert network according to the weights assigned to each expert network by the second gating network to obtain the fourth fusion feature; the second tower network is used to perform a deeper feature extraction on the fourth fusion feature to obtain the fourth probability pCTR4 of downloading the reference application after being seen by the sample user in the historical search results; the third processing unit is used to determine the estimated download rate of the sample user for the reference application in the historical search results according to Proseen1 and pCTR4. For example, the product of Proseen1 and pCTR4 can be determined as the estimated download rate of the sample user for the reference application in the historical search results. At this time, the first gating network, the first processing unit and the first tower network do not perform data processing.

[0096] After obtaining the estimated download rate of the sample users for the reference application in the historical search results, the model parameters of the basic model can be updated according to the difference between the estimated download rate of the sample users for the reference application in the historical search results and the corresponding download label, specifically updating the model parameters of the first gated network, the second gated network, each expert network, the first tower network, the second tower network and the location network.

[0097] Figure 6 FIG. 1 is a flow chart showing a search method according to an exemplary embodiment. Figure 6 As shown, the search method may include the following S201 to S203.

[0098] In S201 , in response to receiving a search request, a target search result matching the search request is obtained.

[0099] In the present disclosure, the target search results include multiple target applications; the search request includes a target search keyword and comes from a target user.

[0100] In S202, the estimated download rate of each target application is determined respectively by using a pre-trained download rate estimation model.

[0101] In the present disclosure, the download rate prediction model is obtained by training based on the download rate prediction model training method provided in the present disclosure. In the actual application stage, the MMoE network of the download rate prediction model can be used to determine the estimated download rate of each target application.

[0102] Specifically, the fourth characteristic information of the target user, the target search keyword, and the fifth characteristic information of the target search keyword can be obtained first; then, for each target application, the sixth characteristic information of the target application can be obtained; based on the fourth characteristic information, the target search keyword, the fifth characteristic information, the target application, and the sixth characteristic information of the target application, a feature vector of the target application is generated; the feature vector of the target application is input into the MMoE network of the above-mentioned download rate estimation model to obtain the estimated download rate of the target application.

[0103] In S203, the target search results are sorted according to the estimated download rate of each target application, and the sorted target search results are displayed.

[0104] Specifically, each target application in the target search result may be sorted in descending order according to the estimated download rate.

[0105] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, the search behavior data of sample users under the search path is obtained, wherein the search path is used to characterize the switching path between search scenarios; then, based on the search behavior data, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model. In this way, the entire user search behavior link can be modeled, so that the search business scenario can be decoupled to achieve systematic modeling, thereby improving the estimation accuracy of the download rate estimation model, so that the top search results are consistent with the user's search needs, improving the accuracy of the output search results and user experience, and thereby improving the download rate of the search engine. In addition, by training the model based on a reinforcement learning mechanism of human feedback, the download rate estimation model can be iteratively optimized, further improving the estimation accuracy of the download rate estimation model.

[0106] Figure 7 FIG. 1 is a block diagram of a download rate prediction model training device according to an exemplary embodiment. Figure 7 As shown, the download rate prediction model training device 300 may include: a first acquisition module 301, configured to obtain search behavior data of sample users under a search path, wherein the search path is used to characterize a switching path between search scenarios; a training module 302, configured to perform model training based on a reinforcement learning mechanism of human feedback according to the search behavior data to obtain a download rate prediction model.

[0107] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, the search behavior data of sample users under the search path is obtained, wherein the search path is used to characterize the switching path between search scenarios; then, based on the search behavior data, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model. In this way, the entire user search behavior link can be modeled, so that the search business scenario can be decoupled to achieve systematic modeling, thereby improving the estimation accuracy of the download rate estimation model, so that the top search results are consistent with the user's search needs, improving the accuracy of the output search results and user experience, and thereby improving the download rate of the search engine. In addition, by training the model based on a reinforcement learning mechanism of human feedback, the download rate estimation model can be iteratively optimized, further improving the estimation accuracy of the download rate estimation model.

[0108] Optionally, the search behavior data includes: the first search sub-behavior data of the sample user from the search results page to the search details page, wherein the first search sub-behavior data includes the historical search results of the search results page and the actual download status of the sample user on the search results page; the second search sub-behavior data of the sample user from the search suggestion page to the search results page and the search details page, wherein the second search sub-behavior data includes the suggested results of the search suggestion page and the actual download status of the sample user on the search suggestion page; the download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after triggering the search.

[0109] Optionally, the training module 302 includes: a first training sub-module, configured to perform supervised fine-tuning training on the basic model according to the first search sub-behavior data to obtain a supervised fine-tuning model; a second training sub-module, configured to fine-tune the supervised fine-tuning model through a reinforcement learning method according to the second search sub-behavior data and the download behavior sequence to obtain a download rate prediction model.

[0110] Optionally, the basic model is a location-bias-aware learning model including a multi-gated hybrid expert network and a location network, and the historical search results and the recommended results are both composed of reference applications; the second training submodule includes: a first acquisition submodule, configured to acquire the first feature information of the sample user, the search keyword corresponding to the search behavior data, and the second feature information of the search keyword; a second acquisition submodule, configured to acquire the third feature information of the reference application for each of the reference applications in the recommended results; a generation submodule, configured to generate a feature vector of the reference application based on the first feature information, the search keyword, the second feature information, the reference application, the third feature information and the download behavior sequence; a determination submodule, configured to determine the download label of the sample user for the reference application based on the actual download situation of the sample user on the search suggestion page; a fine-tuning submodule, configured to fine-tune the multi-gated hybrid expert network of the supervised fine-tuning model through a reinforcement learning method according to each of the feature vectors and each of the download labels to obtain a download rate prediction model.

[0111] Optionally, the multi-gated hybrid expert network includes a first gating network, a second gating network, multiple expert networks, a first tower network, a second tower network, a first processing unit, a second processing unit and a third processing unit; the fine-tuning submodule is configured to perform reinforcement learning training on the multi-gated hybrid expert network of the supervised fine-tuning model by taking the feature vector of the reference application as the input of the first gating network, the second gating network and each of the expert networks, taking the output of each of the expert networks and the output of the first gating network as the input of the first processing unit, taking the output of the first processing unit and the first sub-feature vector as the input of the first tower network, taking the output of each of the expert networks and the output of the second gating network as the input of the second processing unit, taking the output of the second processing unit and the second sub-feature vector as the input of the second tower network, taking the output of the first tower network and the output of the second tower network as the input of the third processing unit, and taking the download label of the sample user for the reference application as the target output of the third processing unit to obtain a download rate estimation model, wherein the first sub-feature vector is a feature vector generated according to the reference application and the third feature information, and the second sub-feature vector is a feature vector generated according to the download behavior sequence.

[0112] Optionally, the search behavior data includes search behavior data of the sample user under two search engines.

[0113] Regarding the download rate prediction model training device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the download rate prediction model training method, and will not be elaborated here.

[0114] Figure 8 FIG. 1 is a block diagram of a search device according to an exemplary embodiment. Figure 8 As shown, the search device 400 includes: a second acquisition module 401, configured to obtain target search results matching the search request in response to receiving a search request, wherein the target search results include multiple target applications; an estimation module 402, configured to determine the estimated download rate of each of the target applications respectively through a pre-trained download rate estimation model, wherein the download rate estimation model is trained based on the above-mentioned download rate estimation model training method provided in the present disclosure; a display module 403, configured to sort the target search results according to the estimated download rate of each of the target applications, and display the sorted target search results.

[0115] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, the search behavior data of sample users under the search path is obtained, wherein the search path is used to characterize the switching path between search scenarios; then, based on the search behavior data, a model is trained based on a reinforcement learning mechanism of human feedback to obtain a download rate estimation model. In this way, the entire user search behavior link can be modeled, so that the search business scenario can be decoupled to achieve systematic modeling, thereby improving the estimation accuracy of the download rate estimation model, so that the top search results are consistent with the user's search needs, improving the accuracy of the output search results and user experience, and thereby improving the download rate of the search engine. In addition, by training the model based on a reinforcement learning mechanism of human feedback, the download rate estimation model can be iteratively optimized, further improving the estimation accuracy of the download rate estimation model.

[0116] Regarding the search device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the relevant search method, and will not be elaborated here.

[0117] The present disclosure also provides an electronic device, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to: implement the steps of the above-mentioned download rate estimation model training method or the above-mentioned search method provided by the present disclosure when executing the executable instructions.

[0118] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the download rate estimation model training method or the steps of the search method provided by the present disclosure are implemented.

[0119] Fig. 9 1 is a block diagram of an apparatus 500 for training a download rate prediction model according to an exemplary embodiment. For example, the apparatus 500 for training a download rate prediction model may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0120] Reference Fig. 9 The device 500 for download rate estimation model training may include one or more of the following components: a first processing component 502, a first memory 504, a first power supply component 506, a first multimedia component 508, a first audio component 510, a first input / output interface 512, a first sensor component 514, and a first communication component 516.

[0121] The first processing component 502 generally controls the overall operation of the apparatus 500 for download rate estimation model training, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The first processing component 502 may include one or more first processors 520 to execute instructions to complete all or part of the steps of the above-mentioned download rate estimation model training method. In addition, the first processing component 502 may include one or more modules to facilitate the interaction between the first processing component 502 and other components. For example, the first processing component 502 may include a multimedia module to facilitate the interaction between the first multimedia component 508 and the first processing component 502.

[0122] The first memory 504 is configured to store various types of data to support the operation of the apparatus 500 for download rate estimation model training. Examples of such data include instructions for any application or method operating on the apparatus 500 for download rate estimation model training, contact data, phone book data, messages, pictures, videos, etc. The first memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0123] The first power supply component 506 provides power to various components of the apparatus for download rate estimation model training 500. The first power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus for download rate estimation model training 500.

[0124] The first multimedia component 508 includes a screen that provides an output interface between the device 500 for download rate estimation model training and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the first multimedia component 508 includes a front camera and / or a rear camera. When the device 500 for download rate estimation model training is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0125] The first audio component 510 is configured to output and / or input audio signals. For example, the first audio component 510 includes a microphone (MIC), and when the device 500 for download rate estimation model training is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the first memory 504 or sent via the first communication component 516. In some embodiments, the first audio component 510 also includes a speaker for outputting an audio signal.

[0126] The first input / output interface 512 provides an interface between the first processing component 502 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0127] The first sensor assembly 514 includes one or more sensors for providing various aspects of status assessment for the apparatus 500 for download rate estimation model training. For example, the first sensor assembly 514 can detect the open / closed state of the apparatus 500 for download rate estimation model training, the relative positioning of components, such as the display and keypad of the apparatus 500 for download rate estimation model training, the first sensor assembly 514 can also detect the position change of the apparatus 500 for download rate estimation model training or a component of the apparatus 500 for download rate estimation model training, the presence or absence of user contact with the apparatus 500 for download rate estimation model training, the orientation or acceleration / deceleration of the apparatus 500 for download rate estimation model training, and the temperature change of the apparatus 500 for download rate estimation model training. The first sensor assembly 514 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The first sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the first sensor component 514 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0128] The first communication component 516 is configured to facilitate wired or wireless communication between the apparatus 500 for download rate estimation model training and other devices. The apparatus 500 for download rate estimation model training can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the first communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the first communication component 516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0129] In an exemplary embodiment, the device 500 for training a download rate prediction model can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned download rate prediction model training method.

[0130] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 504 including instructions, and the instructions can be executed by the first processor 520 of the apparatus 500 for training a download rate prediction model to complete the above-mentioned download rate prediction model training method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0131] In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned download rate prediction model training method when executed by the programmable device.

[0132] Fig.10 8 is a block diagram of a device 800 for searching according to an exemplary embodiment. For example, the device 800 for searching may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0133] Reference Fig.10 , the device 800 for searching may include one or more of the following components: a second processing component 802, a second memory 804, a second power component 806, a second multimedia component 808, a second audio component 810, a second input / output interface 812, a second sensor component 814, and a second communication component 816.

[0134] The second processing component 802 generally controls the overall operation of the device 800 for searching, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The second processing component 802 may include one or more second processors 820 to execute instructions to complete all or part of the steps of the above-mentioned search method. In addition, the second processing component 802 may include one or more modules to facilitate the interaction between the second processing component 802 and other components. For example, the second processing component 802 may include a multimedia module to facilitate the interaction between the second multimedia component 808 and the second processing component 802.

[0135] The second memory 804 is configured to store various types of data to support the operation of the device for searching 800. Examples of such data include instructions for any application or method operating on the device for searching 800, contact data, phone book data, messages, pictures, videos, etc. The second memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0136] The second power supply component 806 provides power to various components of the apparatus 800 for searching. The second power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 800 for searching.

[0137] The second multimedia component 808 includes a screen providing an output interface between the device 800 for searching and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the second multimedia component 808 includes a front camera and / or a rear camera. When the device 800 for searching is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0138] The second audio component 810 is configured to output and / or input audio signals. For example, the second audio component 810 includes a microphone (MIC), and when the device 800 for searching is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal may be further stored in the second memory 804 or sent via the second communication component 816. In some embodiments, the second audio component 810 also includes a speaker for outputting an audio signal.

[0139] The second input / output interface 812 provides an interface between the second processing component 802 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0140] The second sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800 for searching. For example, the second sensor assembly 814 can detect the open / closed state of the device 800 for searching, the relative positioning of components, such as the display and keypad of the device 800 for searching, the second sensor assembly 814 can also detect the position change of the device 800 for searching or a component of the device 800 for searching, the presence or absence of user contact with the device 800 for searching, the orientation or acceleration / deceleration of the device 800 for searching, and the temperature change of the device 800 for searching. The second sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The second sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the second sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0141] The second communication component 816 is configured to facilitate wired or wireless communication between the device 800 for searching and other devices. The device 800 for searching can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the second communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the second communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0142] In an exemplary embodiment, the device 800 for searching can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above-mentioned search method.

[0143] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a second memory 804 including instructions, and the instructions can be executed by the second processor 820 of the apparatus 800 for searching to complete the above-mentioned search method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0144] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for executing the above-mentioned search method when executed by the programmable device.

[0145] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0146] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A download rate prediction model training method, characterized in that: include: Acquire search behavior data of sample users under a search path, wherein the search path is used to characterize a switching path between search scenarios; According to the search behavior data, model training is performed based on a reinforcement learning mechanism of human feedback to obtain a download rate prediction model.

2. The method according to claim 1, characterized in that The search behavior data includes: The first search sub-behavior data of the sample user from the search results page to the search details page, wherein the first search sub-behavior data includes the historical search results of the search results page and the actual download situation of the sample user on the search results page; The second search sub-behavior data of the sample user from the search suggestion page to the search results page and the search details page, wherein the second search sub-behavior data includes the suggested results of the search suggestion page and the actual download situation of the sample user on the search suggestion page; The download behavior sequence of the sample user on the search results page and the search suggestion page within a preset time after the search is triggered.

3. The method according to claim 2, characterized in that The model training is performed based on the search behavior data and the reinforcement learning mechanism based on human feedback to obtain a download rate estimation model, including: According to the first search sub-behavior data, supervised fine-tuning training is performed on the basic model to obtain a supervised fine-tuning model; According to the second search sub-behavior data and the download behavior sequence, the supervised fine-tuning model is fine-tuned by a reinforcement learning method to obtain a download rate prediction model.

4. The method according to claim 3, characterized in that The basic model is a location deviation-aware learning model including a multi-gated hybrid expert network and a location network, and the historical search results and the recommended results are both composed of reference applications; The method of fine-tuning the supervised fine-tuning model by a reinforcement learning method according to the second search sub-behavior data and the download behavior sequence to obtain a download rate estimation model includes: Acquire first characteristic information of the sample user, a search keyword corresponding to the search behavior data, and second characteristic information of the search keyword; For each of the reference applications in the recommendation results, obtaining the third feature information of the reference application; generating a feature vector of the reference application according to the first feature information, the search keyword, the second feature information, the reference application, the third feature information, and the download behavior sequence; determining the download tag of the sample user for the reference application according to the actual download situation of the sample user on the search suggestion page; According to each of the feature vectors and each of the download labels, the multi-gated hybrid expert network of the supervised fine-tuning model is fine-tuned by a reinforcement learning method to obtain a download rate estimation model.

5. The method according to claim 4, characterized in that The multi-gated hybrid expert network includes a first gating network, a second gating network, a plurality of expert networks, a first tower network, a second tower network, a first processing unit, a second processing unit, and a third processing unit; The method of fine-tuning the multi-gated hybrid expert network of the supervised fine-tuning model by a reinforcement learning method according to each of the feature vectors and each of the download labels to obtain a download rate estimation model includes: The multi-gated hybrid expert network of the supervised fine-tuning model is subjected to reinforcement learning training by using the feature vector of the reference application as the input of the first gating network, the second gating network, and each of the expert networks, using the output of each of the expert networks and the output of the first gating network as the input of the first processing unit, using the output of the first processing unit and the first sub-feature vector as the input of the first tower network, using the output of each of the expert networks and the output of the second gating network as the input of the second processing unit, using the output of the second processing unit and the second sub-feature vector as the input of the second tower network, using the output of the first tower network and the output of the second tower network as the input of the third processing unit, and using the download label of the sample user for the reference application as the target output of the third processing unit, so as to obtain a download rate estimation model, wherein the first sub-feature vector is a feature vector generated according to the reference application and the third feature information, and the second sub-feature vector is a feature vector generated according to the download behavior sequence.

6. The method according to any one of claims 1 to 5, characterized in that The search behavior data includes the search behavior data of the sample user under two search engines.

7. A search method, characterized in that: include: In response to receiving a search request, obtaining target search results matching the search request, wherein the target search results include a plurality of target applications; Determine the estimated download rate of each target application respectively by using a pre-trained download rate estimation model, wherein the download rate estimation model is trained based on the download rate estimation model training method according to any one of claims 1 to 6; The target search results are sorted according to the estimated download rate of each target application, and the sorted target search results are displayed.

8. A download rate prediction model training device, characterized in that: include: A first acquisition module is configured to acquire search behavior data of sample users under a search path, wherein the search path is used to represent a switching path between search scenarios; The training module is configured to perform model training based on the search behavior data and a reinforcement learning mechanism based on human feedback to obtain a download rate prediction model.

9. A search device, characterized in that: include: A second acquisition module is configured to, in response to receiving a search request, acquire target search results that match the search request, wherein the target search results include a plurality of target applications; An estimation module is configured to determine the estimated download rate of each of the target applications respectively by using a pre-trained download rate estimation model, wherein the download rate estimation model is trained based on the download rate estimation model training method according to any one of claims 1 to 6; The display module is configured to sort the target search results according to the estimated download rate of each target application and display the sorted target search results.

10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: When the executable instructions are executed, the steps of the method described in any one of claims 1 to 7 are implemented.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.