Model training method and device for eliminating bias, search method and device, and equipment

By introducing satisfaction features into the search ranking model and training it with ranking and bias models, the influence of document genre on the ranking of search documents is eliminated, solving the problem of satisfaction bias in existing technologies and improving the accuracy of search results and user satisfaction.

CN117709431BActive Publication Date: 2026-04-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2023-12-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing search ranking models are biased, resulting in search results that do not meet user needs. In particular, the bias in satisfaction leads to users having less confidence in consuming the remaining documents.

Method used

By acquiring search terms, document genres, and interaction tags from search session data, we extract search term features, document features, and satisfaction features. Using ranking and bias models for training, we eliminate the influence of document genres on the ranking of search documents and improve ranking accuracy.

Benefits of technology

It improved the accuracy of search result ranking and user satisfaction with search results, and reduced the display of documents that were not relevant to users' actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117709431B_ABST
    Figure CN117709431B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a model training method, a search method and device and equipment for eliminating deviation. The document style of each search document is obtained, the satisfaction feature of each search document is obtained according to the document style of each search document and the document style of the search document before the search document, the satisfaction feature of each search document is input into a deviation model, the satisfaction deviation prediction value of each search document is obtained, the influence of the document style on the ranking of the search document is eliminated through the deviation model, the ranking model can learn the influence of the document style of each search document on the ranking of the search document, the ranking accuracy of the search result is improved, and the satisfaction of the user to the search result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data search technology, and in particular to a model training method, search method, apparatus and device for eliminating bias. Background Technology

[0002] With the rapid development of computer technology, video has become a primary medium for people to obtain information and enjoy entertainment in their daily lives. Faced with massive amounts of video content, text-based video search is a commonly used method. Users enter search terms into the search box, and the search engine retrieves relevant search documents based on those terms. A search ranking model then sorts these documents and displays them to the user. Users can consume these search documents, such as clicking, playing, saving, and sharing them. However, existing search ranking models have some biases, resulting in the displayed search documents not always meeting user needs. Summary of the Invention

[0003] This application provides a model training method, search method, apparatus, and device for eliminating bias, which improves the ranking accuracy of search results and enhances user satisfaction with search results.

[0004] In a first aspect, embodiments of this application provide a model training method for eliminating bias, the method comprising:

[0005] Acquire search session data, which includes: search terms, at least one search document corresponding to the search terms, the document type of the search document, and the interactive tags of the search document;

[0006] Based on the search session data, search term features, document features of each search document, and satisfaction features are obtained, wherein the satisfaction features are obtained based on the document genre of each search document and the document genre of previous search documents.

[0007] The search term features and the document features of each search document are input into the ranking model to obtain the predicted interaction probability of each search document.

[0008] The satisfaction features of each search document are input into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0009] The ranking model and the deviation model are trained based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document.

[0010] Secondly, embodiments of this application provide a search method, including:

[0011] Receive a search request, which includes search terms;

[0012] Based on the search request, obtain the search term features and the document features of at least one search document corresponding to the search term;

[0013] The search term features and the document features of the at least one search document are input into the ranking model trained by the method described in the first aspect of this application to obtain the predicted interaction probability of the at least one search document;

[0014] The ranking result of the at least one search document is determined based on the predicted interaction probability of the at least one search document.

[0015] Thirdly, embodiments of this application provide a model training apparatus for eliminating bias, comprising:

[0016] The session acquisition module is used to acquire search session data, which includes: search terms, at least one search document corresponding to the search terms, the document type of the search document, and the interactive tags of the search document;

[0017] The feature acquisition module is used to acquire search term features, document features of each search document, and satisfaction features based on the search session data, wherein the satisfaction features are acquired based on the document genre of each search document and the document genre of previous search documents.

[0018] The prediction module inputs the search term features and the document features of each search document into the ranking model to obtain the predicted interaction probability of each search document.

[0019] The deviation processing module is used to input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0020] The update module is used to train the ranking model and the deviation model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document.

[0021] Fourthly, embodiments of this application provide a search device, comprising:

[0022] A receiving module is used to receive a search request, wherein the search request includes search terms;

[0023] The acquisition module is used to acquire search term features and document features of at least one search document corresponding to the search term based on the search request;

[0024] The ranking module is used to input the search term features and the document features of the at least one search document into the ranking model trained by the bias-free model training device described in the third aspect, to obtain the predicted interaction probability of the at least one search document, and to determine the ranking result of the at least one search document based on the predicted interaction probability of the at least one search document.

[0025] Fifthly, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method as described in the first aspect above.

[0026] In a sixth aspect, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method as described in the second aspect above.

[0027] In a seventh aspect, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the method described in the first aspect above.

[0028] Eighthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the method described in the second aspect above.

[0029] Ninthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first or second aspect above.

[0030] The model training method, search method, apparatus, and device for eliminating bias provided in this application embodiment obtain the document genre of each search document, acquire the satisfaction features of each search document based on the document genre of each search document and the document genres of the search documents preceding it, input the satisfaction features of each search document into the bias model, obtain the satisfaction bias prediction value of each search document, and eliminate the influence of document genre on the ranking of search documents through the bias model, so that the ranking model can learn the influence of the document genre of each search document on the ranking of search documents, improve the ranking accuracy of search results, and enhance user satisfaction with search results. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 A flowchart of the model training method for eliminating bias provided in Embodiment 1 of this application;

[0033] Figure 2 This is a schematic diagram illustrating a scenario for training a ranking model applicable to an embodiment of this application;

[0034] Figure 3 A flowchart of the model training method for eliminating bias provided in Embodiment 2 of this application;

[0035] Figure 4 A flowchart of the model training method for eliminating bias provided in Embodiment 3 of this application;

[0036] Figure 5 This is a schematic diagram illustrating a scenario for multi-objective training of a ranking model applicable to embodiments of this application;

[0037] Figure 6 A flowchart of a search method provided in Embodiment 4 of this application;

[0038] Figure 7 This is a schematic diagram of the structure of the model training device for eliminating bias provided in Embodiment 5 of this application;

[0039] Figure 8 This is a schematic diagram of the structure of the search device provided in Embodiment Six of this application;

[0040] Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment 7 of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0043] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or solution described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0044] It is understood that before using the technical solutions disclosed in the embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application and their authorization obtained in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving a user's active request, a prompt message can be sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation of the technical solution of this application, based on the prompt message. As an optional but non-limiting implementation, the way to send a prompt message to the user in response to receiving a user's active request can be, for example, a pop-up window, in which the prompt message can be presented in text form. Furthermore, the pop-up window can also include a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of this application; other methods that comply with relevant laws and regulations can also be applied to the implementation of this application. It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0045] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more, that is, at least two. "At least one" means one or more.

[0046] In search scenarios, various biases exist, including typical ones such as position bias, exposure bias, selection bias, and popularity bias. These biases cause search results displayed to users to fail to meet their needs.

[0047] In addition to the typical biases mentioned above, there is also a satisfaction bias in search scenarios. The search results list usually includes many documents (docs) or items. The documents included in the search results list are also called search documents, search content, or search resources. These multiple documents are sorted by a ranking model. Users usually consume a document that they are interested in and then do not consume the remaining documents, or a consumption transfer problem occurs, which leads to a lack of confidence in the processing of the remaining documents, i.e., satisfaction bias.

[0048] In reality, the fact that a user consumes a document they are interested in and then stops consuming the remaining documents doesn't necessarily mean the remaining documents are bad or uninteresting. However, ranking models treat these remaining documents as negative samples during training, thus deviating from reality. Consumption transfer refers to a situation where, after consuming a document of interest, a user may see content unrelated to their search intent. However, because the user has already found something of interest, perhaps to kill time or for other purposes, they consume one of the unrelated remaining documents. During training, the ranking model treats this unrelated document as a positive sample, further discrepancies from reality.

[0049] In this application embodiment, the document consumption operation is also referred to as the document interaction operation. This consumption operation or interaction operation includes, but is not limited to, viewing, browsing, playing, saving, clicking, and forwarding the document. This application embodiment can be applied to any existing search scenario, such as a video search scenario, where the documents in the search results may be videos, images, text, etc.

[0050] The existence of satisfaction bias may result in the appearance of some documents that are not relevant to the user's actual needs or the exclusion of documents that the user is interested in after the ranking model has sorted the search results, thus failing to meet the user's search needs.

[0051] To address the aforementioned technical problems, the method in this application introduces satisfaction features based on the document genre of the searched documents during the training of the ranking model. This enables the ranking model to learn the influence of the document genre on the ranking of searched documents, thereby eliminating satisfaction bias, improving the ranking accuracy of search results, and enhancing user search satisfaction.

[0052] The model training method provided in this application embodiment can be applied to large-scale streaming training processes, which are also known as online learning. Large-scale refers to the large amount of data samples used in training.

[0053] Online learning involves continuously feeding new training data into the model, allowing it to learn from and update itself, rather than feeding all training data into the model at once. For example, online learning inputs a batch of training data at a time, and updates the model's parameters after training is complete based on that batch. The number of data points in a batch can be flexibly adjusted according to actual training needs, such as 200, 256, or other quantities.

[0054] The ranking model trained in this application embodiment can be applied to a platform with search functionality, referred to as a search platform. This search platform can be an application software (APP) with search functionality, such as a video player, audio player, reading software, shopping software, etc., or a webpage, mini-program, etc. with search functionality. This application embodiment does not limit this.

[0055] The technical solutions of this application will be described in detail below through some embodiments. The embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0056] Figure 1 This is a flowchart of a model training method for eliminating bias provided in Embodiment 1 of this application. The method in this embodiment can be executed by a training device, which can be a terminal or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop computer, smart voice interaction device, wearable device, etc. The server can be a server cluster or a single server, and can be a cloud server, etc. Figure 1 As shown, the method provided in this embodiment includes the following steps.

[0057] S101. Obtain search session data, which includes: search terms, at least one search document corresponding to the search terms, the document genre of the search document, and the interactive tags of the search document.

[0058] A search session refers to the process by which a user interacts with several search documents returned by the client after entering a search term and initiating a search request. The search session data generated by a search session includes: the search term, at least one search document corresponding to the search term, and the interactive tags of the search document.

[0059] Users can enter search terms by text in the search box, by voice, or by image input. The search terms reflect the user's search intent.

[0060] Optionally, the search session data includes personalized user information entered by the user, including but not limited to user identity (UID), user gender, user age group, and user city. This personalized information is entered or selected by the user with their consent. The user ID is used to uniquely identify a user; it is a number without physical meaning and is used to distinguish different users. For example, regarding user gender, age group, and city, the search platform can provide users with gender options, multiple age group options, and multiple city options. Users select their gender, age group, and city from the options provided by the search platform and confirm.

[0061] The search document can be in various formats, including videos, music, text, and images. The document genre of a search document indicates one or more of the document's type, style, source, and duration, and can be a specific format. For example, in a video search scenario, when a user enters the search term "movie A," the returned search results list might include the complete video of movie A, multiple different clips of movie A, multiple different narration videos of movie A, multiple different behind-the-scenes clips of movie A, multiple different interview videos of movie A, etc. The complete video of movie A is usually a legitimate, official video; in this case, the complete video of movie A can be considered a genre distinct from other videos. Alternatively, the search results list might also include a text collection of movie A, similar to a Baidu Encyclopedia entry; this text collection can be considered a genre distinct from other videos.

[0062] Interaction tags for search documents indicate whether a user has interacted with the document. For example, a click indicates whether the user clicked the document, and a play indicates whether the user played the document. A search document can have one or more interaction tags. For instance, a video can include both play and complete playback tags, with complete playback indicating that the video was played in its entirety.

[0063] S102. Based on the search session data, obtain the search term features, document features of each search document, and satisfaction features. The satisfaction features are obtained based on the document genre of each search document and the document genre of previous search documents.

[0064] The training device acquires search term features based on the search terms, including search term ID, basic terms features, and search intent. Since users can input different types of search terms using different input methods, different acquisition methods are used for different types of search terms.

[0065] For example, when the search term is a text search term, feature extraction can be performed on the text search term to obtain search term features. When the search term is a speech search term, speech recognition is first performed on the speech search term to obtain the corresponding text content, and then feature extraction is performed on the text content to obtain search term features. When the search term is an image search term, image recognition is first performed on the image search term to obtain the corresponding text content, and then feature extraction is performed on the text content to obtain search term features.

[0066] Optionally, when extracting features from text search terms or text content, a neural network model can be used to vectorize the text search terms or text content to obtain search term features corresponding to the search terms. Of course, other existing feature extraction methods can also be used, and this application embodiment does not limit this.

[0067] Optionally, when the search session data includes user-inputted personalized information, the training device obtains user personalized features based on the user personalized information. These user personalized features are used to characterize user interests and hobbies, and content that the user is interested in can be returned based on these features to improve user satisfaction with search results. The training device can use a neural network model to extract user personalized features.

[0068] The training device extracts features from each search document to obtain document features, which include, but are not limited to, document type, document identifier, or other high-level features. The training device can use a neural network model for document feature extraction.

[0069] To eliminate satisfaction bias, this application embodiment introduces satisfaction features. The training device obtains satisfaction features based on the document genre of each search document and the document genres of the previous search documents. Satisfaction feature bias is used to indicate whether the user is satisfied with the search document or the degree of user satisfaction with the search document. The determination of whether the user is satisfied with the search document or the degree of user satisfaction can be determined based on the genre of the search document.

[0070] For example, the satisfaction deviation for each search document is obtained as follows:

[0071] Each search document is traversed according to the display order of at least one search document in the search session data. It is determined whether the document type of the current search document is a preset document type. Based on the determination result, the first satisfaction feature of the current search document is determined. It is also determined whether the search documents before the current search document include search documents of the preset document type. Based on the determination result, the second satisfaction feature of the current search document is determined.

[0072] The search documents included in the search session data are arranged in the order of presentation. The presentation order refers to the order in which the search documents included in the search session data are presented to the user. The training device determines the satisfaction deviation of each search document in turn based on this presentation order.

[0073] Taking video search as an example, the preset document type can be the complete video or text collection summary from the official source mentioned above. This is just an example, and the preset document type can also be other types.

[0074] The training device determines whether the document genre of the currently searched document is a preset document genre. If the document genre is a preset document genre, the value of the first satisfaction feature is set to a first preset value. If the document genre is not a preset document genre, the value of the satisfaction feature is set to a second preset value. For example, the first preset value is 1 and the second preset value is 0.

[0075] Similarly, it is determined whether the search documents preceding the current search document include search documents of the preset document type. If the search documents preceding the current search document include search documents of the preset document type, then the value of the second satisfaction feature is determined to be the first preset value; if the search documents preceding the current search document do not include search documents of the preset document type, then the value of the second satisfaction feature is determined to be the second preset value. For example, the first preset value is 1, and the second preset value is 0.

[0076] The preset document format is usually a document format that the user is satisfied with. When the search results include a search document in the preset document format, the user will not consume any subsequent search documents after interacting with the search document in the preset document format.

[0077] S103. Input the search term features and the document features of each search document into the ranking model to obtain the predicted interaction probability of each search document.

[0078] The ranking model is used to rank searched documents. This ranking model can be a deep neural network (DNN), which is a multi-layer neural network, that is, a neural network with multiple hidden layers.

[0079] The ranking model outputs the predicted interaction probability for each search document, which is the predicted probability that a user will interact with each search document. For example, when the interaction is play, the ranking model outputs the predicted probability that a user will play each search document; when the interaction is click, the ranking model outputs the predicted probability that a user will click each search document.

[0080] If user personalization features are obtained, the search term features, user personalization features, and document features of each search document are input into the ranking model to obtain the predicted interaction probability of each search document.

[0081] S104. Input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0082] The bias model is used to predict the degree of influence of bias features on the interaction of search documents. The bias model can be a shallow neural network, for example, a single-layer neural network.

[0083] Figure 2 This is a schematic diagram illustrating a scenario for training a ranking model applicable to an embodiment of this application, such as... Figure 2 As shown, this scenario includes two models: a ranking model and a bias model. The ranking model is also called the main model or main tower, while the bias model is also called the secondary model or shallow tower. The ranking model takes search term features, user personalization features, and document features for each search document as input, and outputs the predicted interaction logit (Logit) value for the search document. The bias model takes the satisfaction features of the search document as input and outputs the bias logit value for the satisfaction of the search document.

[0084] In model training, the logit value can be understood as the feature output of each layer of the model. Models typically have many layers; the more layers, the better the model's expressive power. Each layer has an activation function, which provides non-linear computation, allowing the model to have more expressive forms. The logit can be understood as the input to the activation function.

[0085] The logit value, after being processed by the activation function, yields a probability between 0 and 1. Similarly, the predicted interaction logit output by the ranking model, after being processed by the activation function, yields the predicted interaction probability. Likewise, the satisfaction deviation logit value output by the bias model, after being processed by the activation function, yields the predicted satisfaction deviation value. For example, this activation function is the sigmoid function.

[0086] It should be noted that the execution order of the above steps S103 and S104 may be as follows: step S103 is executed first and then step S104 is executed; or step S104 is executed first and then step S103 is executed; or steps S103 and S104 are executed simultaneously. This application embodiment does not impose any restrictions on this.

[0087] In real-world scenarios, ranking models exhibit other biases besides satisfaction bias, such as position bias. Accordingly, other bias features for each search document are obtained from search session data. The satisfaction features and other bias features of each search document are combined to obtain the combined bias features of each search document. This combined bias feature is then input into the bias model to obtain the combined bias prediction value for each search document.

[0088] Optionally, the satisfaction features and other bias features of the searched documents can be combined in the following two ways:

[0089] Method 1 involves concatenating the satisfaction features and other bias features of each search document to obtain the combined bias features of each search document.

[0090] The `concat` function is used to concatenate two or more strings or arrays. For example, concatting location features and satisfaction features yields combined bias features.

[0091] Method 2: Perform feature cross-referencing on the satisfaction features and other bias features of each search document to obtain the combined bias features of each search document.

[0092] Feature crosses refer to the process of multiplying two or more features to achieve a nonlinear transformation of the sample space, thereby increasing the nonlinearity of the model. In other words, a nonlinear mapping function f(x) is used to map samples from the original space to the feature space.

[0093] S105. Train the ranking model and the bias model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​for each search document.

[0094] The interaction tag of a search document represents its actual interaction probability. For example, when a search document is clicked, its interaction tag value is 1, and its actual interaction probability is 1. When a search document is not clicked, its interaction tag value is 0, and its actual interaction probability is 0. The loss of the ranking model is calculated based on the interaction tag, predicted interaction probability, and predicted satisfaction deviation value for each search document. The parameters of the ranking model and the deviation model are updated based on the loss of the ranking model until the training conditions are met, at which point model training stops.

[0095] In one exemplary manner, a first loss can be obtained based on the actual interaction probability and the predicted interaction probability of the search document; a second loss is calculated based on the sum of the predicted interaction probability and the predicted satisfaction deviation of the search document, as well as the actual interaction probability; and the loss of the ranking model is calculated based on the first loss and the second loss.

[0096] The loss of a ranking model can be calculated using mean square error (MSR), mean absolute error (MSR), cross entropy loss, focal loss, relative entropy, or exponential loss.

[0097] In this embodiment, the document genre of each search document is obtained, and the satisfaction features of each search document are obtained based on the document genre of each search document and the document genres of the previous search documents. The satisfaction features of each search document are input into the bias model to obtain the satisfaction bias prediction value of each search document. The bias model eliminates the influence of document genre on the ranking of search documents, so that the ranking model can learn the influence of the document genre of each search document on the ranking of search documents, thereby improving the ranking accuracy of search results and increasing user satisfaction with search results.

[0098] Point-wise ranking is a common modeling method used in ranking models. However, point-wise ranking only considers the relationship between search terms and individual search documents, without taking into account the relative order and relationship between multiple documents, which leads to poor ranking performance.

[0099] In a search session, the consumption of documents before and after the searched document is very valuable information. Therefore, in this embodiment, two new objectives are introduced for each searched document: a confidence result for the skip objective (real-skip) and a confidence result for the non-skip objective (real-non-skip).

[0100] The interaction tags of a search document reflect its consumption status. Each search document has two consumption statuses: skipped and non-skipped. Skipped means the search document has not been consumed, while non-skipped means it has been consumed. Whether a search document has been consumed can be determined based on its interaction tags. For example, an interaction tag value of 1 indicates that the search document has been consumed, while an interaction tag value of 0 indicates that the search document has not been consumed.

[0101] Skip and non-skip targets determined based on the interaction tags of a single search document may be erroneous and have low confidence. In this embodiment, skip and non-skip targets determined based on the interaction tags of the current search document and its adjacent and preceding search documents have higher confidence. In this embodiment, skip and non-skip targets determined based on the interaction tags of the current search document and its adjacent and preceding search documents are referred to as real-skip and real-non-skip targets.

[0102] The training device obtains confidence results for skipped targets and non-skipped targets for each search document based on the interaction tags of the adjacent search documents of each search document.

[0103] For example, each search document is traversed according to the display order of at least one search document in the search session data. Based on the interaction tags of the current search document and the interaction tags of its immediate successor search documents, a confidence result for skipped targets of the current search document is obtained. Based on the interaction tags of the current search document and the interaction tags of its immediate successor search documents, a confidence result for non-skipped targets of the current search document is obtained.

[0104] For each search document in the search session data, there is one preceding neighboring search document and one following neighboring search document. For the first search document in the search session data, which has no preceding neighboring search documents, a fixed interaction label for the preceding neighboring search document is assigned when determining the confidence result for the non-skipped target of this search document. Similarly, for the last search document in the search session data, which has no following neighboring search documents, a fixed interaction label for the following neighboring search document is assigned when determining the confidence result for the skipped target of this search document.

[0105] For example, real-skip and real-non-skip can both take values ​​of 0 and 1. If the interaction tag of the current search document indicates that the current document is skipped, and the interaction tag of the next adjacent search document indicates that the next adjacent search document is not skipped, then the real-skip of the current search document is 1; otherwise, the real-skip of the current search document is 0.

[0106] If the interaction tag of the current search document indicates that the current document has not been skipped, and the interaction tag of the preceding adjacent search document indicates that the preceding adjacent search document has been skipped, then the real-non-skip of the current search document is 1; otherwise, the real-non-skip of the current search document is 0.

[0107] The methods described in Examples 2 and 3 below can be used to model based on the newly added real-skip and real-non-skip.

[0108] Figure 3 The flowchart of the model training method for eliminating bias provided in Embodiment 2 of this application is as follows: Figure 3 As shown, the method provided in this embodiment includes the following steps:

[0109] S201. Obtain search session data, which includes: search terms, user-inputted personalized information, at least one search document corresponding to the search terms, the document genre of the search document, and the interactive tags of the search document.

[0110] S202. Based on the search session data, obtain search term features, user personalization features, document features of each search document, and satisfaction features. The satisfaction features are obtained based on the document genre of each search document and the document genre of previous search documents.

[0111] S203. Input the search term features, user personalization features, and document features of each search document into the ranking model to obtain the predicted interaction probability of each search document.

[0112] S204. Input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0113] The specific implementation of steps S201-S204 is described in the relevant description of steps S101-S104 in Embodiment 1, and will not be repeated here.

[0114] S205. Based on the interaction tags of adjacent search documents for each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document.

[0115] Referring to the aforementioned description, it will not be repeated here.

[0116] S206. Calculate the loss of the ranking model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​for each search document.

[0117] The loss of the ranking model can be calculated using existing methods, and this application does not limit this calculation. For example, the loss of the ranking model can be the loss corresponding to the original non-skip objective.

[0118] Optionally, when the ranking model has multiple objectives, it is necessary to calculate the loss for each objective separately, as the loss for each objective is different.

[0119] For example, the loss of the ranking model is calculated using the following formula (1):

[0120] loss=(1-α)×Ll+α×Lr (1)

[0121] Where α is the weight parameter, which is a hyperparameter of the model. α is an adjustable parameter and can be flexibly set according to the model's prediction accuracy; for example, α can be set to 0.5. Ll is the second loss, and Lr is the first loss.

[0122] The first loss is calculated based on the actual interaction probability and the predicted interaction probability of the search document, while the second loss is calculated based on the sum of the predicted interaction probability and the predicted value of the satisfaction deviation, plus the actual interaction probability. The second loss considers the impact of the satisfaction deviation on the model, while the first loss does not. By weighting and summing the first and second losses, the original feature prediction ability can be maintained, and the degree of debiasing can be controlled by adjusting the α parameter.

[0123] S207. Adjust the loss of the ranking model based on the confidence results of skipped targets and non-skipped targets for each search document to obtain the final loss of the ranking model.

[0124] Optionally, the training device determines the label weight of each search document based on the confidence results of skipped targets, the confidence results of non-skipped targets, and the adjusted weights. Based on the label weight of each search document and the loss of the ranking model, the final loss of the ranking model is determined.

[0125] For example, the label weight, label_weight, is calculated according to the following formula:

[0126] label_weight=reweight_scale*(real_skip+real_non_skip);

[0127] Here, `reweight_scale` represents the adjustment of weights and is a hyperparameter of the model. `reweight_scale` is an adjustable parameter; for example, `reweight_scale` can be set to 0.2. `real_skip` represents the confidence result for skipped targets, and `real_non_skip` represents the confidence result for non-skipped targets.

[0128] For example, the final loss is calculated according to the following formula:

[0129] final_loss=model_loss*(1+label_weight);

[0130] Here, `model_loss` represents the loss of the ranking model, and `label_weight` represents the label weight. When the ranking model uses cross-entropy loss, `model_loss` can be the cross-entropy loss for binary classification.

[0131] S208. Update the parameters of the ranking model and the bias model based on the final loss of the ranking model.

[0132] Optionally, the ranking model employs multi-objective training, an existing technique, and the specific training process will not be described in detail in this embodiment. During multi-objective training, the loss corresponding to each objective needs to be calculated separately. Accordingly, the loss corresponding to some objectives can be adjusted based on real-skip and real-non-skip, or the loss corresponding to all objectives can be adjusted.

[0133] In this embodiment, by obtaining the confidence results of skip targets and non-skip targets for each search document based on the interaction tags of adjacent search documents, the loss of the ranking model is adjusted according to these confidence results to obtain the final loss of the ranking model. Adjusting the loss of the ranking model based on the confidence results of skip targets and non-skip targets ensures that the ranking model considers the relevance between search documents in the search session, making the loss of the ranking model more accurate and improving the ranking accuracy of search results, thereby enhancing user search satisfaction.

[0134] Figure 4 The flowchart of the model training method for eliminating bias provided in Embodiment 3 of this application is as follows: Figure 4 As shown, the method provided in this embodiment includes the following steps:

[0135] S301. Obtain search session data, which includes: search terms, user-inputted personalized information, at least one search document corresponding to the search terms, the document genre of the search document, and the interactive tags of the search document.

[0136] S302. Based on the search session data, obtain search term features, user personalization features, document features of each search document, and satisfaction features. The satisfaction features are obtained based on the document genre of each search document and the document genre of previous search documents.

[0137] S303. Input the search term features, user personalization features, and document features of each search document into the ranking model to obtain the predicted interaction probability of each search document.

[0138] S304. Input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0139] The specific implementation of steps S301-S304 is described in the relevant description of steps S101-S104 in Embodiment 1, and will not be repeated here.

[0140] S305. Based on the interaction tags of adjacent search documents for each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document.

[0141] Referring to the aforementioned description, it will not be repeated here.

[0142] S306. Based on the interaction tags of each search document, the confidence results of skipped targets, the confidence results of non-skipped targets, the predicted interaction probability, and the predicted satisfaction deviation value, perform multi-objective training on the ranking model and the deviation model. The multi-objective includes the confidence result objective of skipped targets and the confidence result objective of non-skipped targets.

[0143] In this embodiment, two new objectives, real-skip and real-non-skip, are directly added to the ranking model to enable multi-objective training. Optionally, the ranking model itself can also be a multi-objective model before adding the real-skip and real-non-skip objectives.

[0144] For example, the ranking model itself includes the following three objectives: non-skip, go detail, and finish. If real-skip and real-non-skip are added on top of these three existing objectives, then the ranking model has a total of 5 objectives.

[0145] Multi-objective training is an existing technique, and the specific training process will not be described in detail in this embodiment. During multi-objective training, it is necessary to calculate the loss for each objective separately. Correspondingly, a first label corresponding to the confidence result of the skip objective is added to each search document, and a second label corresponding to the confidence result of the non-skip objective is added to each search document. When calculating the loss for real-skip and real-non-skip objectives, it is necessary to calculate based on the first and second labels corresponding to real-skip and real-non-skip objectives.

[0146] When performing multi-objective training, interference or imitation effects may occur between different objectives. Optionally, a separate bias model can be set up for each objective, so that data is not shared between objectives, data for each objective is isolated, and satisfaction bias for each objective is eliminated separately. This avoids interference or imitation between objectives during multi-objective learning and improves the accuracy of the trained model.

[0147] refer to Figure 5 , Figure 5 This is a schematic diagram of a multi-objective training scenario for a ranking model applicable to an embodiment of this application. In this scenario, a separate bias model is set up for each objective to eliminate satisfaction bias and other biases.

[0148] In this embodiment, by adding real-skip and real-non-skip objectives to the model and employing multi-objective training, the same objective as in Embodiment 2 can be achieved.

[0149] In this embodiment, by obtaining the confidence results of skipped targets and non-skipped targets for each search document based on the interaction tags of adjacent search documents, the ranking model and the bias model are trained using a multi-objective approach. This is done by using the interaction tags, the confidence results of skipped targets, the confidence results of non-skipped targets, the predicted interaction probability, and the predicted satisfaction deviation value for each search document. By adding multi-objective objectives, including the confidence result objective for skipped targets and the confidence result objective for non-skipped targets, the ranking model considers the relevance between search documents in the search session. This makes the loss function of the ranking model more accurate, improving the ranking accuracy of search results and enhancing user search satisfaction.

[0150] After a detailed description of the training process of the ranking model, Embodiment 4 of this application illustrates the application of the ranking model. The ranking models trained in Embodiments 1 to 3 can be applied in search platforms to rank search results. Figure 6 A flowchart of a search method provided in Embodiment 4 of this application is shown below. Figure 6 As shown, the method may include the following steps:

[0151] S401. Receive a search request, which includes search terms.

[0152] The search term can be a text search term, a voice search term, or an image search term entered by the user through a text box. Optionally, the search request may also include personalized information entered by the user.

[0153] S402. Obtain the search term features and the document features of at least one search document corresponding to the search term based on the search request.

[0154] Databases contain a large number of documents, and it is impossible to display all documents to the user during each search. The documents displayed to the user need to be filtered. Therefore, search platforms include not only ranking models, but also retrieval systems or retrieval models. The retrieval model is used to filter candidate search documents from the database to be displayed to the user based on the search request, and then the ranking model ranks the candidate search documents.

[0155] Optionally, the search platform also obtains user personalized features based on the search request. In one approach, when the search request does not include user personalized information, the retrieval model first obtains the user personalized information based on the search request, and then filters and displays search documents to the user from the database based on the user personalized information and search terms. When the search request includes user personalized information, the search platform directly filters and displays search documents to the user from the database based on the user personalized information and search terms.

[0156] When the search request includes user personalized information, the user personalized features are obtained based on the user personalized information, and the document features of the search document are obtained based on the search document obtained by the retrieval model. The specific acquisition method is described in Example 1, and will not be repeated here.

[0157] S403. Input the search term features and the document features of at least one search document into the ranking model to obtain the predicted interaction probability of at least one search document.

[0158] The ranking model is trained using any of the methods described in Examples 1 to 3. The ranking model outputs the predicted interaction probability for each search document based on search term features and document features of at least one search document. Optionally, if the search platform has acquired user-personalized features, the search term features, user-personalized features, and document features of at least one search document are input together into the ranking model to obtain the predicted interaction probability for each search document.

[0159] S404. Determine the ranking result of at least one search document based on the predicted interaction probability of at least one search document.

[0160] Search documents are sorted based on the predicted probability of interaction, and the results are displayed to the user. Search documents with a higher predicted probability of interaction are more likely to be consumed by the user, so they are prioritized and placed at the top of the search results to improve the accuracy of the ranking.

[0161] It's understandable that search results can be displayed in a split-screen format based on the number of documents in the user's sorted results, allowing users to browse more documents by flipping through pages.

[0162] The search method in this embodiment eliminates satisfaction bias during the training process of the ranking model, thereby eliminating the satisfaction bias in the trained ranking model. The ranking results output by the ranking model are more accurate, thus meeting the user's search needs and improving the user's search satisfaction.

[0163] To facilitate better implementation of the model training method for eliminating bias in the embodiments of this application, the embodiments of this application also provide a model training apparatus for eliminating bias. Figure 7 This is a schematic diagram of the structure of the model training device for eliminating bias provided in Embodiment 5 of this application, as shown below. Figure 7 As shown, the bias-eliminating model training device 100 may include:

[0164] The session acquisition module 11 is used to acquire search session data, which includes: search terms, at least one search document corresponding to the search terms, the document type of the search document, and the interactive tags of the search document;

[0165] The feature acquisition module 12 is used to acquire search term features, document features of each search document, and satisfaction features based on the search session data, wherein the satisfaction features are acquired based on the document genre of each search document and the document genre of previous search documents.

[0166] Prediction module 13 inputs the search term features and document features of each search document into the ranking model to obtain the predicted interaction probability of each search document;

[0167] The deviation processing module 14 is used to input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document.

[0168] The update module 15 is used to train the ranking model and the deviation model based on the interaction tags, predicted interaction probabilities and satisfaction deviation prediction values ​​of each search document.

[0169] In one alternative implementation, the satisfaction deviation of each search document is obtained as follows:

[0170] According to the display order of the at least one search document in the search session data, each search document is traversed to determine whether the document genre of the current search document is a preset document genre, and the first satisfaction feature of the current search document is determined according to the determination result.

[0171] Determine whether the search documents preceding the current search document include search documents of the preset document genre, and determine the second satisfaction feature of the current search document based on the determination result.

[0172] In one alternative implementation, the feature acquisition module 12 is further configured to:

[0173] Other deviation features of each search document are obtained based on the search session data;

[0174] The satisfaction features and other bias features of each search document are combined to obtain the combined bias features of each search document;

[0175] The deviation processing module 14 is also used for:

[0176] The combined bias features are input into the bias model to obtain the combined bias prediction value for each search document.

[0177] In one alternative implementation, the feature acquisition module 12 is specifically used for:

[0178] The satisfaction features and other bias features of each search document are concatted to obtain the combined bias features of each search document.

[0179] In one alternative implementation, the feature acquisition module 12 is specifically used for:

[0180] For each search document, the satisfaction feature and other bias features are cross-referenced to obtain the combined bias features of each search document.

[0181] In an alternative implementation, the session acquisition module 11 is further configured to:

[0182] Based on the interaction tags of the adjacent search documents of each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document;

[0183] In one alternative implementation, the update module 15 is specifically used for:

[0184] Based on the interaction tags of each search document, the confidence results of skipped targets, the confidence results of non-skipped targets, the predicted interaction probability, and the predicted satisfaction deviation value, the ranking model and the deviation model are trained using a multi-objective method, wherein the multi-objective method includes the confidence result objective of skipped targets and the confidence result objective of non-skipped targets.

[0185] In an alternative implementation, the session acquisition module 11 is further configured to:

[0186] Based on the interaction tags of the adjacent search documents of each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document;

[0187] The update module 15 is specifically used for:

[0188] The loss of the ranking model is calculated based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​for each search document.

[0189] Based on the confidence results of skipped targets and non-skipped targets for each search document, the loss of the ranking model is adjusted to obtain the final loss of the ranking model.

[0190] The parameters of the ranking model and the parameters of the bias model are updated based on the final loss of the ranking model.

[0191] In one alternative implementation, the session acquisition module 11 is specifically used for:

[0192] According to the display order of the at least one search document in the search session data, each search document is traversed, and the confidence result of the skip target of the current search document is obtained according to the interaction tag of the current search document and the interaction tag of the next adjacent search document of the current search document.

[0193] Based on the interaction tags of the current search document and the interaction tags of the preceding adjacent search documents of the current search document, obtain the confidence result of the non-skipped target of the current search document.

[0194] In one alternative implementation, the session acquisition module 11 is specifically used for:

[0195] If the interaction tag of the current search document indicates that the current document has been skipped, and the interaction tag of the next adjacent search document indicates that the next adjacent search document has not been skipped, then the confidence result of the skip target of the current search document is 1; otherwise, the confidence result of the skip target of the current search document is 0.

[0196] The step of obtaining the confidence result of the non-skipped target of the current search document based on the interaction tags of the current search document and the interaction tags of the preceding adjacent search documents of the current search document includes:

[0197] If the interaction tag of the current search document indicates that the current document has not been skipped, and the interaction tag of the preceding adjacent search document indicates that the preceding adjacent search document has been skipped, then the confidence result of the non-skipped target of the current search document is 1; otherwise, the confidence result of the non-skipped target of the current search document is 0.

[0198] In one alternative implementation, the update module 15 is specifically used for:

[0199] The tag weight of each search document is determined based on the confidence results of skipped targets, the confidence results of non-skipped targets, and the adjusted weights.

[0200] The final loss of the ranking model is determined based on the tag weight of each search document and the loss of the ranking model.

[0201] In one alternative implementation, the update module 15 is specifically used for:

[0202] The label weight, label_weight, is calculated using the following formula:

[0203] label_weight=reweight_scale*(real_skip+real_non_skip);

[0204] Wherein, reweight_scale represents the adjusted weight, real_skip represents the confidence result of the skipped target, and real_non_skip represents the confidence result of the non-skipped target;

[0205] The final loss, final_loss, is calculated according to the following formula:

[0206] final_loss=model_loss*(1+label_weight);

[0207] Where model_loss represents the loss of the ranking model, and label_weight represents the label weight.

[0208] In one optional implementation, the interactive tags of the search document include multiple tags, and the ranking model is a multi-objective model.

[0209] In one alternative implementation, the search session data further includes user-inputted personalized information, and the feature acquisition module 12 is further configured to:

[0210] User personalization features are obtained based on the user personalization information in the search session data;

[0211] The prediction module 13 is specifically used for:

[0212] The search term features, the user personalization features, and the document features of each search document are input into the ranking model to obtain the predicted interaction probability of each search document.

[0213] The apparatus of this embodiment can be used to execute any of the methods described in Embodiments 1 to 3 above. The specific implementation method is described in the method embodiment, and will not be repeated here.

[0214] Figure 8 This is a schematic diagram of the structure of the search device provided in Embodiment Six of this application, as shown below. Figure 8 As shown, the search device 200 may include:

[0215] Receiving module 21 is used to receive a search request, wherein the search request includes search terms;

[0216] The acquisition module 22 is used to acquire the search term features and the document features of at least one search document corresponding to the search term according to the search request;

[0217] The sorting module 23 is used to input the search term features and the document features of the at least one search document into the model training device 100 for eliminating bias provided in Embodiment 5 to train the sorting model, obtain the predicted interaction probability of the at least one search document, and determine the sorting result of the at least one search document based on the predicted interaction probability of the at least one search document.

[0218] In an alternative implementation, the acquisition module 22 is further configured to:

[0219] Obtain user-personalized characteristics based on user-input personalized information;

[0220] The sorting module 23 is specifically used for:

[0221] The search term features, the user personalization features, and the document features of the at least one search document are input into the ranking model to obtain the predicted interaction probability of the at least one search document.

[0222] The apparatus in this embodiment can be used to execute the search method provided in Embodiment 4 above. The specific implementation method is described in the method embodiment, and will not be repeated here.

[0223] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here.

[0224] The apparatuses 100 and 200 of this application embodiment have been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that these functional modules can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the methods disclosed in this application embodiment can be directly manifested as execution by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0225] This application also provides an electronic device. Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment 7 of this application, as shown below. Figure 9 As shown, the electronic device 300 may include:

[0226] The system includes a memory 31 and a processor 32. The memory 31 stores computer programs and transfers the program code to the processor 32. In other words, the processor 32 can retrieve and run the computer programs from the memory 31 to implement the bias-eliminating model training or search method provided in the embodiments of this application.

[0227] For example, the processor 32 can be used to execute the bias-eliminating model training method or search method provided in the above method embodiments according to the instructions in the computer program.

[0228] In some embodiments of this application, the processor 32 may include, but is not limited to: a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0229] In some embodiments of this application, the memory 31 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0230] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0231] like Figure 9 As shown, the electronic device 300 may further include a transceiver 33, which can be connected to the processor 32 or the memory 31.

[0232] The processor 32 can control the transceiver 33 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 33 may include a transmitter and a receiver. The transceiver 33 may further include antennas, and the number of antennas may be one or more.

[0233] Understandable, although Figure 9As not shown in the diagram, the electronic device 300 may also include a camera module, a Wi-Fi module, a positioning module, a Bluetooth module, a display, a controller, etc., which will not be described in detail here.

[0234] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0235] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods described in the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the bias-eliminating model training method or search method provided in the above-described method embodiments.

[0236] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the corresponding processes of the bias-eliminating model training method and search method provided in the method embodiments. For simplicity, these will not be elaborated further here.

[0237] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0238] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0239] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A model training method for eliminating bias, characterized in that, include: Acquire search session data, which includes: search terms, at least one search document corresponding to the search terms, the document type of the search document, and the interactive tags of the search document; Based on the search session data, search term features, document features of each search document, and satisfaction features are obtained, wherein the satisfaction features are obtained based on the document genre of each search document and the document genre of previous search documents. The search term features and the document features of each search document are input into the ranking model to obtain the predicted interaction probability of each search document. The satisfaction features of each search document are input into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document. The ranking model and the deviation model are trained based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document. The satisfaction deviation for each search document is obtained in the following way: According to the display order of the at least one search document in the search session data, each search document is traversed to determine whether the document genre of the current search document is a preset document genre, and the first satisfaction feature of the current search document is determined according to the determination result. Determine whether the search documents preceding the current search document include search documents of the preset document genre, and determine the second satisfaction feature of the current search document based on the determination result.

2. The method according to claim 1, characterized in that, Before inputting the satisfaction features of each search document into the bias model, the method further includes: Other deviation features of each search document are obtained based on the search session data; The satisfaction features and other bias features of each search document are combined to obtain the combined bias features of each search document; The step of inputting the satisfaction features of each search document into the bias model includes: The combined bias features are input into the bias model to obtain the combined bias prediction value for each search document.

3. The method according to claim 2, characterized in that, The combination of satisfaction features and other bias features for each search document to obtain combined bias features for each search document includes: The satisfaction features and other bias features of each search document are concatted to obtain the combined bias features of each search document.

4. The method according to claim 2, characterized in that, The combination of satisfaction features and other bias features for each search document to obtain combined bias features for each search document includes: For each search document, the satisfaction feature and other bias features are cross-referenced to obtain the combined bias features of each search document.

5. The method according to claim 1, characterized in that, Also includes: Based on the interaction tags of the adjacent search documents of each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document; The step of training the ranking model and the deviation model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document includes: Based on the interaction tags of each search document, the confidence results of skipped targets, the confidence results of non-skipped targets, the predicted interaction probability, and the predicted satisfaction deviation value, the ranking model and the deviation model are trained using a multi-objective method, wherein the multi-objective method includes the confidence result objective of skipped targets and the confidence result objective of non-skipped targets.

6. The method according to claim 1, characterized in that, Also includes: Based on the interaction tags of the adjacent search documents of each search document, obtain the confidence results of skipped targets and non-skipped targets for each search document; The step of training the ranking model and the deviation model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document includes: The loss of the ranking model is calculated based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​for each search document. Based on the confidence results of skipped targets and non-skipped targets for each search document, the loss of the ranking model is adjusted to obtain the final loss of the ranking model. The parameters of the ranking model and the parameters of the bias model are updated based on the final loss of the ranking model.

7. The method according to claim 5 or 6, characterized in that, The step of obtaining the confidence results of skipped targets and non-skipped targets for each search document based on the interaction tags of adjacent search documents of each search document includes: According to the display order of the at least one search document in the search session data, each search document is traversed, and the confidence result of the skip target of the current search document is obtained according to the interaction tag of the current search document and the interaction tag of the next adjacent search document of the current search document. Based on the interaction tags of the current search document and the interaction tags of the preceding adjacent search documents of the current search document, obtain the confidence result of the non-skipped target of the current search document.

8. The method according to claim 7, characterized in that, The step of obtaining the confidence result of the non-skipped target of the current search document based on the interaction tags of the current search document and the interaction tags of the preceding adjacent search documents of the current search document includes: If the interaction tag of the current search document indicates that the current search document has been skipped, and the interaction tag of the next adjacent search document indicates that the next adjacent search document has not been skipped, then the confidence result of the skip target of the current search document is 1; otherwise, the confidence result of the skip target of the current search document is 0. The step of obtaining the confidence result of the non-skipped target of the current search document based on the interaction tags of the current search document and the interaction tags of the preceding adjacent search documents of the current search document includes: If the interaction tag of the current search document indicates that the current search document has not been skipped, and the interaction tag of the preceding adjacent search document indicates that the preceding adjacent search document has been skipped, then the confidence result of the non-skipped target of the current search document is 1; otherwise, the confidence result of the non-skipped target of the current search document is 0.

9. The method according to claim 6, characterized in that, The step of adjusting the loss of the ranking model based on the confidence results of skipped targets and non-skipped targets for each search document to obtain the final loss of the ranking model includes: The tag weight of each search document is determined based on the confidence results of skipped targets, the confidence results of non-skipped targets, and the adjusted weights. The final loss of the ranking model is determined based on the tag weight of each search document and the loss of the ranking model.

10. The method according to claim 9, characterized in that, The step of determining the tag weight of each search document based on the confidence results of skipped targets, the confidence results of non-skipped targets, and the adjusted weights includes: The label weight, label_weight, is calculated using the following formula: label_weight = reweight_scale * (real_skip + real_non_skip); Wherein, reweight_scale represents the adjusted weight, real_skip represents the confidence result of the skipped target, and real_non_skip represents the confidence result of the non-skipped target; The step of determining the final loss of the ranking model based on the tag weight of each search document and the loss of the ranking model includes: The final loss, final_loss, is calculated according to the following formula: final_loss =model_loss * (1 + label_weight); Where model_loss represents the loss of the ranking model, and label_weight represents the label weight.

11. The method according to any one of claims 1-6, characterized in that, The interactive tags for the searched documents include multiple tags, and the ranking model is a multi-objective model.

12. The method according to any one of claims 1-6, characterized in that, The search session data also includes user-inputted personalized information, and the method further includes: User personalization features are obtained based on the user personalization information in the search session data; The step of inputting the search term features and the document features of each search document into the ranking model to obtain the predicted interaction probability of each search document includes: The search term features, the user personalization features, and the document features of each search document are input into the ranking model to obtain the predicted interaction probability of each search document.

13. A search method, characterized in that, include: Receive a search request, which includes search terms; Based on the search request, obtain the search term features and the document features of at least one search document corresponding to the search term; The search term features and the document features of the at least one search document are input into the ranking model trained according to any one of claims 1-12 to obtain the predicted interaction probability of the at least one search document; The ranking result of the at least one search document is determined based on the predicted interaction probability of the at least one search document.

14. A model training device for eliminating bias, characterized in that, include: The session acquisition module is used to acquire search session data, which includes: search terms, at least one search document corresponding to the search terms, the document type of the search document, and the interactive tags of the search document; The feature acquisition module is used to acquire search term features, document features of each search document, and satisfaction features based on the search session data, wherein the satisfaction features are acquired based on the document genre of each search document and the document genre of previous search documents. The prediction module inputs the search term features and the document features of each search document into the ranking model to obtain the predicted interaction probability of each search document. The deviation processing module is used to input the satisfaction features of each search document into the deviation model to obtain the satisfaction deviation prediction value of each search document. The satisfaction deviation prediction value is used to characterize the degree of influence of the document genre of the search document on the interaction of the search document. The update module is used to train the ranking model and the deviation model based on the interaction tags, predicted interaction probabilities, and predicted satisfaction deviation values ​​of each search document. The satisfaction deviation for each search document is obtained in the following way: According to the display order of the at least one search document in the search session data, each search document is traversed to determine whether the document genre of the current search document is a preset document genre, and the first satisfaction feature of the current search document is determined according to the determination result. Determine whether the search documents preceding the current search document include search documents of the preset document genre, and determine the second satisfaction feature of the current search document based on the determination result.

15. A search device, characterized in that, include: A receiving module is used to receive a search request, wherein the search request includes search terms; The acquisition module is used to acquire search term features and document features of at least one search document corresponding to the search term based on the search request; The ranking module is used to input the search term features and the document features of the at least one search document into the ranking model trained by the model training device for eliminating bias as described in claim 14, to obtain the predicted interaction probability of the at least one search document, and to determine the ranking result of the at least one search document based on the predicted interaction probability of the at least one search document.

16. An electronic device, characterized in that, include: A processor and a memory, the memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the method of any one of claims 1 to 13.

17. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Leveraging usage data of an online resource when estimating future user interaction with the online resource

    CN108536721A

  • Document retrieval method and device, electronic equipment and medium

    CN117056460A