Object search method and device

By constructing candidate data pairs and using object analysis model based on conversion factors, the problem of not being able to take into account both correlation and conversion rate in the prior art is solved, and a more accurate and user-friendly search result display is achieved.

CN114942972BActive Publication Date: 2025-05-09ALIBABA (CHENGDU) SOFTWARE & TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210380226.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-05-09
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

The prior art cannot take into account the relevance and conversion rate of objects in information search, resulting in the impact of user experience and search accuracy.

Method used

By obtaining the target search information and candidate attribute information of candidate objects, a candidate data pair is constructed and inputted into the trained object analysis model. The training sample set based on transformation factors is used to comprehensively analyze the correlation and conversion rate to determine the display order of candidate objects.

Benefits of technology

It improves the accuracy and user experience of search results, balances relevance and conversion rate, and enhances the display effect of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114942972B_ABST
    Figure CN114942972B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an object search method and device, wherein the object search method includes: obtaining a training sample set based on at least one conversion factor sampling in advance, training an object analysis model, and when a user inputs target search information for searching, a corresponding candidate data pair can be constructed based on the target search information and the candidate attribute information of at least one candidate object, and the candidate data pair is input into the trained object analysis model to obtain the relevance score of the candidate data pair, and the display order of at least one candidate object can be determined based on the relevance score of the candidate data pair and displayed to the user, thereby feeding back the search results to the user. In this way, at least one conversion factor can indicate the conversion rate of different dimensions, and the object analysis model incorporates the conversion rate information during training. The trained object analysis model balances the conversion rate and relevance, improves the accuracy of the search results, and ensures the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of information search technology, and more particularly to an object search method. One or more embodiments of this specification also relate to an object search device, a computing device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of computer technology and Internet technology, searching for information through accessing the Internet has become a common means of obtaining information. Users can enter search terms in the search box provided by the search engine to query information and obtain object information that meets their needs.

[0003] In the prior art, the search terms input by the user are often matched with the titles of the objects to be displayed, the matching degree between the objects to be displayed and the search terms is determined, and then the display order of the objects to be displayed is determined, and the search results of the objects are fed back to the user.

[0004] However, the core goals of search are user experience and conversion rate. User experience is to prioritize highly relevant objects to users as much as possible, that is, to ensure that the search results displayed to users are relevant to the search terms entered by users as much as possible. Conversion rate is to display objects with high clicks and high conversion rates in the search results as early as possible. Although the above method takes into account the correlation between each object to be displayed and the search terms, it cannot take into account the conversion rate of each object to be displayed, resulting in an inability to balance the relevance and conversion rate in the search results, affecting the accuracy of the search and further affecting the user experience. Summary of the invention

[0005] In view of this, an embodiment of the present specification provides an object search method. One or more embodiments of the present specification also relate to an object search device, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.

[0006] According to a first aspect of an embodiment of this specification, there is provided an object search method, including:

[0007] Acquire target search information and candidate attribute information of at least one candidate object;

[0008] According to the target search information and the candidate attribute information, a corresponding candidate data pair is constructed;

[0009] Inputting the candidate data pairs into an object analysis model to obtain a relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor;

[0010] According to the relevance scores of the candidate data pairs, a display order of at least one candidate object is determined and displayed.

[0011] Optionally, the candidate attribute information of the candidate object includes object category information, industry information and attribute keywords;

[0012] According to the target search information and the candidate attribute information, the corresponding candidate data pairs are constructed, including:

[0013] According to preset splicing rules, the target search information is spliced ​​with the object category information, industry information, and attribute keywords of the first candidate object to obtain a candidate data pair corresponding to the target search information and the first candidate object, wherein the first candidate object is any one of the at least one candidate object.

[0014] Optionally, the object analysis model includes a feature analysis layer and an output layer;

[0015] The candidate data pairs are input into the object analysis model to obtain the relevance scores of the candidate data pairs, including:

[0016] Inputting the first candidate data pair into the feature analysis layer of the object analysis model to obtain the encoding features of each character in the first candidate data pair, wherein the first candidate data pair is a candidate data pair constructed based on the candidate attribute information of any candidate object;

[0017] The target coding feature of the first candidate data pair is determined by fusing the coding feature of the first character in the first candidate data pair, the average coding feature of each character, and the maximum coding feature;

[0018] The target coding features are input into the output layer of the object analysis model to obtain the relevance score of the first candidate data pair.

[0019] Optionally, determining a display order of at least one candidate object and displaying the at least one candidate object according to the relevance score of the candidate data pair includes:

[0020] According to the relevance scores of each candidate data pair, at least one candidate object is divided into a first set value of gears;

[0021] Sort the candidate objects in each gear based on the set sorting rules;

[0022] At least one candidate object is displayed based on the sorting result and the gear sequence of the first set number of gears.

[0023] Optionally, determining a display order of at least one candidate object and displaying the at least one candidate object according to the relevance score of the candidate data pair includes:

[0024] According to the category information of each candidate data pair, at least one candidate object is divided into a second set value of categories;

[0025] Sort the candidates within each category based on their relevance scores;

[0026] At least one candidate object is displayed based on the sorting result and the display order of the second set value categories.

[0027] Optionally, the object analysis model is trained by the following method:

[0028] Acquire user search data within a preset time, and construct a training sample set according to the user search data, wherein the training sample set includes positive samples and negative samples, and the positive samples and the negative samples are obtained by sampling based on at least one conversion factor;

[0029] Input the training sample set into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set;

[0030] Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated through the preset boundary loss function, the model parameters of the initial analysis model are adjusted according to the loss value, and the operation steps of obtaining user search data within the preset time period are returned until the training stop condition is reached, and the trained object analysis model is obtained.

[0031] Optionally, the user search data includes sample search information and corresponding sample search results, and the sample search results include at least one sample object;

[0032] Construct a training sample set based on user search data, including:

[0033] Obtaining a sample search result corresponding to the sample search information, and determining a first sample object corresponding to at least one conversion factor in the sample search result, and a second sample object that does not have at least one conversion factor;

[0034] A training sample set is constructed according to the sample search information, the first sample object and the second sample object.

[0035] Optionally, constructing a training sample set according to the sample search information, the first sample object and the second sample object includes:

[0036] Determine a reference sample object whose object conversion rate is less than a conversion rate threshold in the first sample object, and obtain reference sample attribute information of the reference sample object;

[0037] Acquire first sample attribute information of sample objects other than the reference sample object in the first sample objects, and second sample attribute information of the second sample objects;

[0038] A positive sample is constructed according to the sample search information and the first sample attribute information, and a negative sample is constructed according to the sample search information and the second sample attribute information and the reference sample attribute information, respectively, to obtain a training sample set.

[0039] Optionally, after constructing a positive sample according to the sample search information and the first sample attribute information, the method further includes:

[0040] Determine a sampling weight corresponding to a conversion factor of a target sample object, wherein the target sample object is any sample object in the first sample object except the reference sample object;

[0041] According to the sampling weights, the positive samples constructed based on the target sample objects are resampled to obtain expanded positive samples.

[0042] Optionally, after obtaining the target search information and before obtaining the candidate attribute information of at least one candidate object, the method further includes:

[0043] Get the object title of at least one object to be displayed;

[0044] At least one candidate object whose object title is related to the target search information is determined from the objects to be displayed.

[0045] According to a second aspect of an embodiment of this specification, there is provided an object search device, including:

[0046] An acquisition module, configured to acquire target search information and candidate attribute information of at least one candidate object;

[0047] A construction module, configured to construct corresponding candidate data pairs according to target search information and candidate attribute information;

[0048] an obtaining module configured to input the candidate data pair into an object analysis model to obtain a relevance score of the candidate data pair, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor;

[0049] The determination module is configured to determine and display the display order of at least one candidate object according to the relevance score of the candidate data pair.

[0050] According to a third aspect of an embodiment of this specification, a computing device is provided, including:

[0051] Memory and processor;

[0052] The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the following method:

[0053] Acquire target search information and candidate attribute information of at least one candidate object;

[0054] According to the target search information and the candidate attribute information, a corresponding candidate data pair is constructed;

[0055] Inputting the candidate data pairs into an object analysis model to obtain a relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor;

[0056] According to the relevance scores of the candidate data pairs, a display order of at least one candidate object is determined and displayed.

[0057] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of any object search method are implemented.

[0058] An embodiment of the present specification provides an object search method, which can obtain target search information and candidate attribute information of at least one candidate object; construct corresponding candidate data pairs according to the target search information and the candidate attribute information; input the candidate data pairs into an object analysis model to obtain the relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained based on at least one conversion factor sampling; and determine the display order of at least one candidate object and display it according to the relevance score of the candidate data pairs. In this case, a training sample set can be obtained in advance based on at least one conversion factor sampling, and an object analysis model can be obtained by training based on the training sample set. When a user inputs target search information for searching, a corresponding candidate data pair can be constructed based on the target search information and the candidate attribute information of at least one candidate object, that is, the candidate data pair corresponding to a candidate object includes the target search information and the candidate attribute information of the candidate object. Each candidate data pair is input into the trained object analysis model, and the relevance score of each candidate data pair can be obtained. The relevance score can indicate the degree of relevance between the candidate attribute information and the target search information in the candidate data pair. Based on the relevance score of each candidate data pair, the display order of at least one candidate object can be determined and displayed to the user, thereby feeding back the search results to the user. In this way, at least one conversion factor can indicate the conversion rate of different dimensions, and the object analysis model incorporates the conversion rate information during training, that is, when the object analysis model learns the correlation between the search information and the object attribute information in the training sample set, it takes into account the conversion rate information of the object. Therefore, the trained object analysis model can comprehensively analyze the correlation between the target search information and the candidate attribute information in the input candidate data, as well as the conversion rate of the candidate object, balance the conversion rate and correlation, improve the accuracy of the search results, and ensure the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flow chart of an object search method provided by an embodiment of this specification;

[0060] Figure 2a It is a schematic diagram of a data processing process of an object analysis model provided by an embodiment of this specification;

[0061] Figure 2b is a schematic diagram of a data processing process of another object analysis model provided by an embodiment of this specification;

[0062] Figure 3 This is a flowchart of an object search method applied in an e-commerce scenario provided by an embodiment of this specification;

[0063] Figure 4 is a schematic diagram of the structure of an object search device provided by an embodiment of this specification;

[0064] Figure 5 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0065] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0066] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0067] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0068] First, the terms involved in one or more embodiments of this specification are explained.

[0069] Relevance: A measure of whether the search terms of users on the search end are relevant to the object content. Offline measurement tables include acc (accuracy), auc (receiver operating characteristic curve), etc. On e-commerce search websites, relevance mainly reflects whether the products returned by the website are similar to the user's search intention. Among them, auc (area under curve) is defined as the area under the ROC curve and the coordinate axis. The value of this area will not be greater than 1. Since the ROC curve is generally above the straight line y=x, the value range of auc is between 0.5 and 1. The closer auc is to 1.0, the higher the authenticity of the detection method. When it is equal to 0.5, the authenticity is the lowest and there is no application value. Auc is a performance indicator to measure the quality of the learner.

[0070] Pre-trained model: A deep semantic model based on the Transformer structure, which is pre-trained using large-scale corpus and text self-supervision. It can accurately model text semantics, and then fine-tune it on downstream tasks to obtain good classification, sorting and other effects. Common open source pre-trained models include BERT, RoBERT, DeBERT, etc.

[0071] EMB representation: An encoding feature representation method of the NLP semantic model that converts human-perceivable information such as text and images into digital representations that can be understood by computers, thereby enabling computers to recognize and apply human intentions.

[0072] It should be noted that the core goals of search are user experience and conversion rate. User experience is to prioritize highly relevant objects to users as much as possible, that is, to ensure that the search results displayed to users are relevant to the search terms entered by users as much as possible, while conversion rate is to display high-click, high-conversion objects in the search results as early as possible. In order to ensure user experience and reduce the exposure of some objects with high conversion rates but not very relevant to search terms, the modeling of correlation has rarely considered conversion rate, resulting in a situation where conversion rate ranking and correlation squeeze each other. In addition, insufficient attention has been paid to features that invisiblely affect conversion effects and user experience, such as differences in user preferences, industry differences, and prices, which have failed to significantly improve object conversion rates and have also affected the user's search experience to a certain extent.

[0073] For example, when a user searches for "XX building blocks", some other brands of building blocks that are cheaper and have higher conversion rates will be ranked at the front in order to improve the conversion rate. Although this can bring clicks from some users, it ignores the real needs of users searching for the object. Considering only the conversion rate of the object will cause data bias to become more and more serious in the long run. Later searches for "XX building blocks" will basically result in products from other brands, which reduces the accuracy of search results and greatly affects user experience. Relevance considers the relevance between the search term and the object, and only displays "XX building blocks" to users. Although this guarantees user experience, it ignores price, industry differences, user preferences, etc., which often leads to a very low object conversion rate.

[0074] Currently, after a user enters a search term in the search box provided by the search engine, the final display order needs to be determined based on the relevance between the search term entered by the user and the text information of the object (most commonly the object title) during the search ranking stage. There are two common methods for calculating the relevance between the search term and the text information of the object title: keyword relevance and semantic relevance based on deep learning.

[0075] Keyword relevance: Calculates the number of word matches between the search term and the object title, such as subject words, keywords, attribute words, etc. Different words have different relevance weights. This type of relevance matching has high accuracy, but cannot handle synonyms or misspellings.

[0076] Semantic relevance based on deep learning: The search terms and object titles are represented as word vectors, and the final ranking is determined by calculating the similarity score between the two. This method can effectively alleviate the problem of keyword relevance and solve the OOV problem (Out-Of-Vocabulary problem, that is, the problem of exceeding the preset vocabulary). Among them, there are usually two ways to calculate the similarity score: 1) Represent the search terms and object titles as word vectors respectively, and then calculate the cosine similarity between the two; 2) After the search terms and object titles are segmented, they are input into the same network model to let the model learn a score that represents the similarity between the two. The first method is intuitive to calculate, but the accuracy is low. The accuracy of the second method depends on the performance of the pre-trained model of the deep model used. A common method of precise ranking is to combine keyword relevance with semantic relevance based on deep learning. Currently, more and more projects are trying to use only semantic relevance, which can also bring better search results. However, most of the semantic relevance based on deep learning currently adopts the general pre-trained language model BERT structure based on Transformer, which has the following main characteristics: the searched object titles and search terms are usually not complete sentences, but the superposition of a large number of keywords, which makes it difficult for the BERT model to understand the text semantics; the relevance focuses on the text features of the search terms and object titles, and pays insufficient attention to the features that invisiblely affect the conversion effect and user experience, such as user preference differences, industry differences, and prices, and fails to significantly improve the object conversion rate; the modeling method is single, which can distinguish whether it is related or not, but the hierarchical expression of the relevance is insufficient, and it is difficult to differentiate the relevance between the search terms and objects.

[0077] In view of the above situation, the embodiments of this specification integrate conversion rate and relevance from the two aspects of data construction and model design, and provide an object search method: ① At the data construction level, semantically related search information-object attribute information pairs are used as candidate data pairs, and when obtaining the training sample set, different samples are resampled and weighted accordingly according to conversion factors such as clicks, inquiries, orders, and payments; ② In terms of feature fusion, in addition to considering the title content of the object, category information, industry information, and object attribute keywords are also integrated to expand the object richness information; ③ In terms of model design, multi-level fusion of model representation is performed, and the encoding features of search information and object attribute information realize the fusion of the encoding features of the first character, the average encoding features of each character, and the maximum encoding features to form new encoding features, so as to realize the modeling of information of different dimensions; ④ Loss function optimization, innovative use of L-softmax (Large-Margin softmaxloss) to learn the intra-class compactness and inter-class separability of positive and negative samples, so that the model can distinguish objects at different semantic levels more clearly. Based on the optimization of the above four dimensions, the object search method provided in the embodiment of this specification can increase the AUC by 2.9% in the evaluation data of search relevance, and the conversion rate by 1.08%. On the basis of ensuring the improvement of user experience, it greatly increases the conversion effect of the object and balances the relevance and conversion rate in the search results.

[0078] In this specification, an object search method is provided. This specification also relates to an object search device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0079] Figure 1 A flowchart of an object search method provided according to an embodiment of the present specification is shown, including steps 102 to 108.

[0080] Step 102: Obtain target search information and candidate attribute information of at least one candidate object.

[0081] It should be noted that the target search information refers to the search criteria entered by the user in the search box of the search engine. The target search information is used to search for the corresponding object and display it to the user for selection. The target search information can be a complete sentence or an independent word. In addition, the candidate object refers to an object that can be displayed to the user for selection. The object can be a commodity, document, information, etc.; the candidate attribute information of the candidate object refers to information that can represent the characteristics of the candidate object, such as object category information, industry information, and attribute keywords.

[0082] In practical applications, the input information in the search box of the search engine can be read and used as the target search information input by the user. In addition, each candidate object can have a corresponding description file, which stores various information of the candidate object, so the required candidate attribute information can be obtained from the description file of the candidate object.

[0083] In an optional implementation of this embodiment, the candidate object may be an object determined by the search engine after roughly sorting various objects based on the target search information input by the user, that is, after obtaining the target search information and before obtaining the candidate attribute information of at least one candidate object, the following may also be included:

[0084] Get the object title of at least one object to be displayed;

[0085] At least one candidate object whose object title is related to the target search information is determined from the objects to be displayed.

[0086] It should be noted that the search process is generally divided into several stages such as recall, rough sorting, and fine sorting. The object search method provided in the embodiments of this specification is mainly used in the fine sorting stage, that is, the candidate objects determined by recall and rough sorting are further sorted to obtain search results and display them to the user. That is, the subsequent determination of relevance based on the object analysis model mainly acts on the fine sorting stage, and the relevance is calculated for the candidate objects after recall and rough sorting. Generally speaking, the number of candidate objects will not be too large, such as generally not more than 4,000.

[0087] In actual applications, after the user inputs the target search information, the various objects included in the search engine can be matched and recalled according to the categories corresponding to the target search information (such as tops, pants, jackets), the user's attribute information (such as regions, historical preferences), etc., and the corresponding objects to be displayed can be recalled from thousands of objects, and then the objects to be displayed can be roughly sorted, and a set number of objects to be displayed can be selected as candidate objects for subsequent fine sorting. Specifically, based on the correlation between the object title of the object to be displayed and the target search information, each object to be displayed can be roughly sorted to screen out the corresponding candidate objects, that is, when at least one candidate object with an object title related to the target search information is determined from each object to be displayed, the similarity between the object title of the object to be displayed and the target search information can be determined for each object to be displayed in each object to be displayed, and then the objects to be displayed with a similarity higher than a similarity threshold in each object to be displayed are selected as candidate objects, or a set number of objects to be displayed with a high similarity in each object to be displayed are selected as candidate objects.

[0088] In the embodiments of the present specification, the objects to be displayed are various objects included in the search engine. Based on the object title of at least one object to be displayed, candidate objects related to the target search information input by the user can be preliminarily screened out. Subsequently, the candidate objects obtained by the rough screening can be further refined to obtain the final search results, thereby improving the accuracy of the search results.

[0089] Step 104: construct corresponding candidate data pairs according to the target search information and the candidate attribute information.

[0090] It should be noted that the target search information is the search condition input by the user, and the candidate attribute information is the feature information of the candidate object. Therefore, in order to analyze the correlation between the target search information and the candidate attribute information, corresponding candidate data pairs can be constructed based on the target search information and the candidate attribute information. Subsequently, the constructed candidate data pairs can be analyzed to determine the correlation between each candidate object and the target search information, and the final search results can be fed back to the user.

[0091] In an optional implementation of this embodiment, the candidate attribute information of the candidate object may include object category information, industry information and attribute keywords. At this time, the corresponding candidate data pair is constructed according to the target search information and the candidate attribute information. The specific implementation process may be as follows:

[0092] According to preset splicing rules, the target search information is spliced ​​with the object category information, industry information, and attribute keywords of the first candidate object to obtain a candidate data pair corresponding to the target search information and the first candidate object, wherein the first candidate object is any one of the at least one candidate object.

[0093] Specifically, object category information can indicate the category to which the candidate object belongs, industry information can indicate the industry to which the candidate object belongs, and attribute keywords can indicate the detailed features of the candidate object. For example, in an e-commerce scenario, the candidate object is a commodity, and the object category information can be furniture, stationery, electrical appliances, etc., and the industry information can be the beauty industry, film and television industry, manufacturing industry, etc., and the attribute keywords can be size, color, purpose, performance, etc.; in a document search scenario, the candidate object is a document, and the object category information can be text, table, image, etc., and the industry information can be the electronics industry, communications industry, education industry, etc., and the attribute keywords can be abstract, size, number of pages, author, creation time, etc. In addition, in addition to object category information, industry information and attribute keywords, the candidate attribute information of the candidate object can also include object titles, and can also include other information that can characterize the candidate object, such as object value, object image, and other information.

[0094] It should be noted that the preset splicing rules may refer to pre-set rules for splicing search conditions and various information of candidate objects. The preset splicing rules may be the arrangement order of each piece of information. For example, the preset splicing rules may be to splice the target search information with the object category information, industry information, and attribute keywords of the first candidate object in sequence.

[0095] In practical applications, for the first candidate object, the target search information is sequentially spliced ​​with the object category information, industry information, and attribute keywords of the first candidate object to obtain a candidate data pair corresponding to the target search information and the first candidate object. Each candidate object can be used as the first candidate object, and the candidate attribute information of the first candidate object is spliced ​​with the target search information to obtain a candidate data pair corresponding to the first candidate object and the target search information; that is, each candidate object can construct a corresponding candidate data pair, and the number of candidate data pairs is the same as the number of candidate objects.

[0096] In the embodiments of the present specification, various information such as object category information, industry information, and attribute keywords that can characterize the characteristics of candidate objects can be integrated into the target search information, thereby expanding the information richness of the candidate objects, and display and add information such as object category information, industry information, and attribute keywords for the candidate objects. When analyzing candidate data pairs to determine the ranking of candidate objects, various information such as category differences, industry differences, and attribute differences that can characterize the characteristics of candidate objects can be integrated to fit the user's attention to different dimensions of candidate objects, thereby improving the accuracy of search results.

[0097] Step 106: Input the candidate data pairs into the object analysis model to obtain the relevance scores of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor.

[0098] Specifically, the object analysis model is a pre-trained model that can analyze the correlation between target search information and candidate attribute information in the input candidate data pairs. The object analysis model can be a Transformer structure, or it can also be a model result such as BERT, RoBERT, DeBERT, etc.

[0099] In addition, the correlation score can represent the correlation between the target search information and the candidate attribute information in the candidate data pair, that is, the correlation between the corresponding candidate object and the target search information. The correlation score can be a value between 0-1. The higher the correlation score, the higher the correlation between the corresponding candidate object and the target search information. The lower the correlation score, the lower the correlation between the corresponding candidate object and the target search information.

[0100] It should be noted that the conversion factor refers to various operational parameters that can characterize the object conversion rate of candidate objects. The object conversion rate refers to the frequency with which candidate objects are finally acquired by users under the search conditions of the target search information. For example, after displaying various candidate objects to the user, the user's clicks, inquiries, orders, settlements and other operations on a candidate object can all characterize whether the user wants to acquire the candidate object, that is, clicks, inquiries, orders, settlements and other operations on a candidate object may affect the conversion rate of the candidate object, that is, the conversion factor may include operational parameters such as clicks, inquiries, orders and / or settlements.

[0101] In an embodiment of the present specification, at least one conversion factor can indicate conversion rates of different dimensions, and the object analysis model incorporates conversion rate information during training, that is, when the object analysis model learns the correlation between search information and object attribute information in the training sample set, it takes into account the conversion rate information of the object. Therefore, the trained object analysis model can comprehensively analyze the correlation between the target search information and the candidate attribute information in the input candidate data, as well as the conversion rate of the candidate object, thereby balancing the conversion rate and correlation, improving the accuracy of the search results, and ensuring the user experience.

[0102] Furthermore, the word features of each word in the target search information input by the user are different, and the degree of influence on the search results is different. Therefore, NER (Named Entity Recognition) technology can be used to feature multiple words to obtain word feature tags of the target search information. The word feature tags can indicate the degree of influence of each word on the search results. The word feature tags and the constructed candidate data pairs can be input into the object analysis model together, and the information of the feature tag dimensions of each word in the target search information is added. When the object analysis model analyzes the candidate data pairs, it can combine the feature tags of each word in the target search information and pay different attention to different words in the target search information, thereby determining the final relevance score and improving the analysis accuracy of the object analysis model.

[0103] For example, the target search information input by the user is "red dress for women wedding". Through the NER technology, each word in the target search information is labeled, and it can be identified that the subject word is "dress", the color of concern is "red", and the usage scenario is "wedding". These words are relatively important and have a greater impact on the search results. Therefore, these words can be labeled "1" so that the object analysis model pays more attention to these words; and for words with little meaning such as "for" and "women", these words can be labeled "0" so that the object analysis model pays less attention to these words. In this way, the word feature label of the target search information can be obtained as "1,1,0,0,1". The word feature label and the constructed candidate data pairs are input together into the object analysis model, so that the object analysis model pays more attention to words such as "red", "dress", and "wedding".

[0104] In an optional implementation of this embodiment, the object analysis model includes a feature analysis layer and an output layer, and can fuse the encoding features of the candidate data pairs in different dimensions, that is, input the candidate data pairs into the object analysis model to obtain the correlation scores of the candidate data pairs. The specific implementation process can be as follows:

[0105] Inputting the first candidate data pair into the feature analysis layer of the object analysis model to obtain the encoding features of each character in the first candidate data pair, wherein the first candidate data pair is a candidate data pair constructed based on the candidate attribute information of any candidate object;

[0106] The target coding feature of the first candidate data pair is determined by fusing the coding feature of the first character in the first candidate data pair, the average coding feature of each character, and the maximum coding feature;

[0107] The target coding features are input into the output layer of the object analysis model to obtain the relevance score of the first candidate data pair.

[0108] It should be noted that by inputting the first candidate data pair into the feature analysis layer of the object analysis model, the encoding features of each character in the first candidate data pair, namely, the EMB representation, can be obtained, and then the encoding features of the first character (CLS EMB), the average encoding features of each character (AVG EMB) and the maximum encoding features of each character (MAXEMB) can be determined, and then the encoding features of the first character, the average encoding features of each character and the maximum encoding features are fused to obtain the target encoding features of the first candidate data pair, and then the target encoding features of the first candidate data pair are input into the output layer of the object analysis model for analysis, so as to obtain the relevance score of the first candidate data pair.

[0109] Among them, the feature analysis layer of the object analysis model can be at least two layers, and each feature analysis layer will output the coding features of the first character. The coding features of the first character output by the last two layers can be spliced ​​to obtain the coding features of the first character in the final first candidate data pair, and merged with the average coding features and the maximum coding features, similar to the residual strategy, to prevent information loss. In addition, the average coding features of each character can use the coding features of each character in the first candidate data pair to expand the information of the first candidate data pair and improve the accuracy of the final model analysis results.

[0110] Furthermore, in the fusion of deep learning models, using the maximum coding feature is actually saving the maximum value of each dimension. Because in subsequent calculations, the larger the value, the greater the impact on the result. It is generally believed that noise interferes with existing data and generally does not lose features very obviously, so the impact on the maximum value will not be great. Therefore, using the maximum coding feature is to record the most obvious features, which helps to eliminate noise interference.

[0111] In practical applications, when the coding features of the first character in the first candidate data pair, the average coding features of each character, and the maximum coding features are integrated to determine the target coding features of the first candidate data pair, the coding features of the first character, the average coding features of each character, and the maximum coding features can be concatenated in sequence, or corresponding weight coefficients can be set for the coding features of the first character, the average coding features of each character, and the maximum coding features, and then the coding features of the first character, the average coding features of each character, and the maximum coding features can be concatenated in sequence based on the weight coefficients.

[0112] For example, Figure 2a is a schematic diagram of a data processing process of an object analysis model provided by an embodiment of this specification, such as Figure 2a As shown, currently the first candidate data pair can be input into the feature analysis layer in the object analysis model to obtain the encoding features of each character in the first candidate data pair, and then only the encoding feature CLS of the first character is input into the output layer for analysis, and finally the corresponding correlation score is output. Only the features of the first character dimension are considered, and the output layer can analyze the relatively single features, resulting in poor accuracy of the object analysis model in analyzing the correlation. Therefore, the embodiment of this specification provides another processing method, Figure 2b is a schematic diagram of a data processing process of another object analysis model provided by an embodiment of this specification, such as Figure 2bAs shown, the first candidate data pair can be input into the feature analysis layer in the object analysis model to obtain the coding features of each character in the first candidate data pair, and then the coding feature CLS of the first character in the first candidate data pair, the average coding feature AVG and the maximum coding feature MAX of each character are fused to determine the target coding feature of the first candidate data pair, and the target coding feature is input into the output layer for analysis, and finally the corresponding correlation score is output.

[0113] In the embodiments of this specification, the coding features of the first character in the first candidate data pair, the average coding features of each character and the maximum coding features can be fused to determine the target coding features of the first candidate data pair, and then the target coding features are input into the output layer of the object analysis model to obtain the relevance score of the first candidate data pair. In this way, the multi-level coding features of the model can be fused to realize modeling of information of different dimensions and improve the accuracy of the final model analysis results.

[0114] In an optional implementation of this embodiment, the object analysis model is obtained by training using the following method:

[0115] Acquire user search data within a preset time, and construct a training sample set according to the user search data, wherein the training sample set includes positive samples and negative samples, and the positive samples and the negative samples are obtained by sampling based on at least one conversion factor;

[0116] Input the training sample set into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set;

[0117] Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated through the preset boundary loss function, the model parameters of the initial analysis model are adjusted according to the loss value, and the operation steps of obtaining user search data within the preset time period are returned until the training stop condition is reached, and the trained object analysis model is obtained.

[0118] Specifically, statistics can be collected on user searches and conversion operations on search results within a period of time (i.e., within a preset time), that is, a training sample set can be obtained based on at least one conversion factor sampling, and the constructed training sample set can be input into the initial analysis model to train the initial analysis model. The initial analysis model can be a pre-trained model obtained, such as Transformer, BERT, RoBERT, DeBERT, etc.

[0119] In practical applications, when the initial analysis model is trained based on the constructed training sample set, the training sample set is input into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set. Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated by a preset boundary loss function. The loss value can represent the gap between the true value and the predicted value. Based on the loss value, the model parameters of the initial analysis model can be adjusted in reverse, and user search data can continue to be obtained to construct a training sample set to train the initial model until the training stop condition is reached to obtain a trained object analysis model.

[0120] It should be noted that the preset boundary loss function can be L-softmax. One of the fundamental purposes of using L-softmax is to be responsible for the distribution of correlation data. It only cares about whether it can be correctly classified. In fact, correlation should consider the sorting problem, that is, even if they are both positive samples, they should also be distinguished. Since the training samples in the constructed training sample set only have positive samples and negative samples, that is, there are only two categories of 0 and 1, L-softmax in metric learning can be used to distinguish the score differences between positive samples.

[0121] In specific implementation, the loss value of the initial analysis model can be calculated by the preset boundary loss function through the following formula (1) and formula (2):

[0122]

[0123]

[0124] Among them, L i represents the loss value of the initial analysis model; w represents the weight vector of the fully connected layer in the initial analysis model, and x represents the encoding vector output by the previous layer of the fully connected layer, that is, the input vector of the fully connected layer, so ||w j || ||x i ||cos(θ j ) represents the inner product between the input vector and the weight vector of the fully connected layer, where i represents the element position of the input vector of the fully connected layer, i.e. x i represents the i-th element in the input vector of the fully connected layer, and j represents the position of the weight vector element of the fully connected layer, that is, w j represents the jth element of the weight vector of the fully connected layer, and θ represents the angle between the two vectors; represents the output vector of the fully connected layer, Represents the inner product between the input vector and the output vector of the fully connected layer.

[0125] In addition, m is a preset integer closely related to the classification boundary. m determines the strength of the cumulative category close to the predicted score. As m increases, the classification boundary increases, and the learning goal becomes more and more difficult. D(θ) is a preset monotonically decreasing function. To be equal to

[0126] In the embodiments of this specification, L-softmax can be used to learn the intra-class compactness and inter-class separability of positive and negative samples, so that the model can distinguish objects at different semantic levels more clearly and avoid overfitting.

[0127] In an optional implementation of this embodiment, the user search data includes sample search information and corresponding sample search results, and the sample search results include at least one sample object; the training sample set is constructed according to the user search data, and the specific implementation process can be as follows:

[0128] Obtaining a sample search result corresponding to the sample search information, and determining a first sample object corresponding to at least one conversion factor in the sample search result, and a second sample object that does not have at least one conversion factor;

[0129] A training sample set is constructed according to the sample search information, the first sample object and the second sample object.

[0130] It should be noted that the first sample object is a sample object corresponding to at least one conversion factor in the sample search results, that is, the first sample object is a sample object on which conversion factor-related operations exist, that is, the first sample object has been subjected to conversion factor-related operations by the user, and may be more in line with user needs; and the second sample object is a sample object for which at least one conversion factor does not exist in the sample search results, that is, the second sample object is an object displayed to the user, but the user has not performed any operations corresponding to the conversion factor, that is, no conversion has been performed, which may be quite different from user needs.

[0131] In the embodiments of the present specification, a training sample set can be constructed based on the sample search information, the first sample object and the second sample object. The constructed training sample set includes both samples on which conversion factor-related operations have been performed by the user and samples on which no operations corresponding to the conversion factor have been performed. The constructed training sample set has rich sample types, thereby improving the accuracy and efficiency of model training.

[0132] In an optional implementation of this embodiment, in addition to samples that have been exposed but have not performed operations related to the conversion factor as negative samples, samples that have been highly exposed multiple times but have very low conversion rates can also be added as negative samples, that is, a training sample set is constructed according to the sample search information, the first sample object, and the second sample object. The specific implementation process can be as follows:

[0133] Determine a reference sample object whose object conversion rate is less than a conversion rate threshold in the first sample object, and obtain reference sample attribute information of the reference sample object;

[0134] Acquire first sample attribute information of sample objects other than the reference sample object in the first sample objects, and second sample attribute information of the second sample objects;

[0135] A positive sample is constructed according to the sample search information and the first sample attribute information, and a negative sample is constructed according to the sample search information and the second sample attribute information and the reference sample attribute information, respectively, to obtain a training sample set.

[0136] It should be noted that positive samples can be constructed based on sample objects that have conversion factor-related operations within a period of time (i.e., within a preset time), ensuring that only sample objects that have conversion factor-related operations can be used as positive samples, and negative samples can be constructed based on sample objects that have no conversion factor-related operations within a period of time (i.e., within a preset time), that is, samples that have been exposed but have not performed conversion factor-related operations as negative samples.

[0137] In an optional implementation of this embodiment, the reference sample objects whose object conversion rate is less than the conversion rate threshold in the first sample objects may be first determined, the reference sample attribute information of the reference sample objects may be obtained, and negative samples may be constructed according to the sample search information, the second sample attribute information, and the reference sample attribute information. In this way, in addition to using samples that have been exposed but have not performed operations related to the conversion factor as negative samples, sample objects with low conversion rates (less than the conversion rate threshold) in the first sample objects may also be supplemented as negative samples, effectively reducing the impact of objects that are ranked high but not highly relevant due to sorting bias.

[0138] In an optional implementation of this embodiment, different samples may be resampled and weighted accordingly according to conversion factors such as clicks, inquiries, orders, and settlements. That is, after constructing a positive sample according to the sample search information and the first sample attribute information, the following may also be included:

[0139] Determine a sampling weight corresponding to a conversion factor of a target sample object, wherein the target sample object is any sample object in the first sample object except the reference sample object;

[0140] According to the sampling weights, the positive samples constructed based on the target sample objects are resampled to obtain expanded positive samples.

[0141] It should be noted that when users search for objects through search engines, after the search engines display the corresponding search results, users browse most of the objects, but fewer objects perform the conversion operations corresponding to the conversion factors, and the number of objects corresponding to the conversion factors of different dimensions is also different. The depth of the conversion operations corresponding to the conversion factors is deeper, and the fewer the corresponding objects are, that is, the fewer the objects corresponding to the operations with higher conversion rates are. However, the objects on which users perform the conversion operations corresponding to the conversion factors are the positive samples for training the initial model, and their reference significance is greater than the negative samples corresponding to the objects that users only browse, but the number of negative samples is much larger than the positive samples, so the sampling weight corresponding to the conversion factor of the target sample object can be determined, and based on the corresponding sampling weight, the positive samples constructed by the target sample object are resampled to obtain expanded positive samples, thereby increasing the number of positive samples in the training sample set and ensuring the model training results.

[0142] In practical applications, different conversion factors correspond to different conversion depths, and thus different weight coefficients can be adopted. The sampling weight corresponding to the conversion factor with a higher conversion depth can be larger. In specific implementation, the positive samples constructed based on the target sample object are resampled according to the sampling weights. When obtaining the expanded positive samples, the positive samples constructed based on the target sample object can be repeated the number of times corresponding to the sampling weights to obtain the corresponding number of expanded positive samples.

[0143] As an example, for user-oriented projects, the user's click behavior can be divided into the following according to the depth: click on the list page (click) - click the supplier communication button (c-click) - initiate an inquiry to the supplier (feedback) - place an order for the selected object (order) - pay for the ordered object (pay). Multiple sampling is based on the judgment that the deeper the user's click depth, the greater the user's interest in the object, and the more relevant the object information is to the target search information entered by the user. The depth temperature is resampled multiple times. Specifically, the positive samples of order / pay (target search information-candidate attribute information data pair) can be resampled 5 times (that is, the positive samples of order / pay are repeated 5 times), feedback can be resampled 10 times (that is, the positive samples corresponding to feedback are repeated 10 times), and clicks and c-clicks with many clicks but no deep inquiry are sampled 1 and 2 times; feedback, order, and pay samples are more relevant, but the number is small, so multiple sampling can make the model pay more attention to this batch of sample objects with high confidence.

[0144] In the embodiments of the present specification, the number of conversion factors such as clicks, inquiries, orders, and settlements of each positive sample can be counted as a measure of the conversion rate of the positive sample. The conversion rate is used as a resampling indicator, and multiple sampling is performed on positive samples with deep conversion behaviors such as inquiries, orders, and calculations. This effectively eliminates the impact of objects with high exposure and low conversion rates on the correlation. At the same time, positive samples with truly high correlation and high conversion rates are given more attention, thereby improving the accuracy of the model's correlation analysis.

[0145] Step 108: Determine the display order of at least one candidate object and display it according to the relevance score of the candidate data pair.

[0146] It should be noted that the relevance score of the candidate data pair can represent the correlation between the target search information and the candidate attribute information in the candidate data pair, that is, the correlation between the corresponding candidate object and the target search information. Therefore, based on the relevance score of the candidate data pair, the display order of at least one candidate object can be determined and displayed to the user, and the user can select the target object he or she needs based on the displayed candidate objects.

[0147] In an optional implementation of this embodiment, the display order of at least one candidate object is determined and displayed according to the relevance score of the candidate data pair. The specific implementation process may be as follows:

[0148] According to the relevance scores of each candidate data pair, at least one candidate object is divided into a first set value of gears;

[0149] Sort the candidate objects in each gear based on the set sorting rules;

[0150] At least one candidate object is displayed based on the sorting result and the gear sequence of the first set number of gears.

[0151] It should be noted that the set sorting rule may be a pre-set sorting rule that specifies the sorting of candidate objects in the same gear, such as sorting prices from low to high in the same gear.

[0152] For example, each candidate object can be divided into three levels: correlation score greater than 0.6, between 0.3 and 0.6, and less than 0.3, and the price in each level is sorted from low to high.

[0153] In the embodiments of the present specification, at least one candidate object can be divided into a first set numerical number of gears according to the relevance scores of each candidate data pair, and then the candidate objects in each gear are sorted based on the set sorting rules. At least one candidate object is displayed based on the sorting result and the gear order of the first set numerical number of gears, that is, each candidate object is divided into different gears according to the relevance score, and a secondary sorting is performed within each gear, which improves the sorting flexibility of each candidate object. The sorting rules within the gear can be customized, which is highly flexible and can adapt to more scenarios.

[0154] In an optional implementation of this embodiment, the display order of at least one candidate object is determined and displayed according to the relevance score of the candidate data pair. The specific implementation process may be as follows:

[0155] According to the category information of each candidate data pair, at least one candidate object is divided into a second set value of categories;

[0156] Sort the candidates within each category based on their relevance scores;

[0157] At least one candidate object is displayed based on the sorting result and the display order of the second set value categories.

[0158] It should be noted that at least one candidate object can be divided into a second set number of categories based on the category information of each candidate data pair, and the candidate objects in each category can be sorted based on the relevance score, and at least one candidate object can be displayed based on the sorting result and the display order of the second set number of categories. In this way, each candidate object can be sorted in combination with the relevance score and category, and can be displayed based on the category, which is convenient for users to view and improves the user experience.

[0159] An embodiment of the present specification provides an object search method, which can obtain target search information and candidate attribute information of at least one candidate object; construct corresponding candidate data pairs according to the target search information and the candidate attribute information; input the candidate data pairs into an object analysis model to obtain the relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained based on at least one conversion factor sampling; and determine the display order of at least one candidate object and display it according to the relevance score of the candidate data pairs. In this case, a training sample set can be obtained in advance based on at least one conversion factor sampling, and an object analysis model can be obtained by training based on the training sample set. When a user inputs target search information for searching, a corresponding candidate data pair can be constructed based on the target search information and the candidate attribute information of at least one candidate object, that is, the candidate data pair corresponding to a candidate object includes the target search information and the candidate attribute information of the candidate object. Each candidate data pair is input into the trained object analysis model, and the relevance score of each candidate data pair can be obtained. The relevance score can indicate the degree of relevance between the candidate attribute information and the target search information in the candidate data pair. Based on the relevance score of each candidate data pair, the display order of at least one candidate object can be determined and displayed to the user, thereby feeding back the search results to the user. In this way, at least one conversion factor can indicate the conversion rate of different dimensions, and the object analysis model incorporates the conversion rate information during training, that is, when the object analysis model learns the correlation between the search information and the object attribute information in the training sample set, it takes into account the conversion rate information of the object. Therefore, the trained object analysis model can comprehensively analyze the correlation between the target search information and the candidate attribute information in the input candidate data, as well as the conversion rate of the candidate object, balance the conversion rate and correlation, improve the accuracy of the search results, and ensure the user experience.

[0160] The following combination Figure 3 , taking the application of the commodity search method provided in this specification in the e-commerce scenario as an example, the commodity search method is further explained. Figure 3 A flowchart of a commodity search method applied to an e-commerce scenario provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0161] Step 302: Obtain target search information input by the user in the search box of the search engine, and determine at least one candidate product obtained by recall and rough sorting based on the target search information.

[0162] Step 304: According to the preset splicing rules, the target search information is spliced ​​with the product category information, industry information, and attribute keywords of each candidate product to obtain candidate data pairs corresponding to the target search information and each candidate product.

[0163] Step 306: Input the candidate data pair to the feature analysis layer of the commodity analysis model to obtain the encoding features of each character in the candidate data pair.

[0164] Step 308: The coding feature of the first character in the candidate data pair, the average coding feature of each character, and the maximum coding feature are integrated to determine the target coding feature of the candidate data pair.

[0165] Step 310: Input the target coding features into the output layer of the commodity analysis model to obtain the correlation scores of the candidate data pairs.

[0166] In an optional implementation of this embodiment, the commodity analysis model is trained by the following method:

[0167] Acquire user search data within a preset time, and construct a training sample set according to the user search data, wherein the training sample set includes positive samples and negative samples, and the positive samples and the negative samples are obtained by sampling based on at least one conversion factor;

[0168] Input the training sample set into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set;

[0169] Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated through the preset boundary loss function, the model parameters of the initial analysis model are adjusted according to the loss value, and the operation steps of obtaining user search data within the preset time period are returned until the training stop condition is reached to obtain the trained product analysis model.

[0170] In an optional implementation of this embodiment, the user search data includes sample search information and corresponding sample search results, and the sample search results include at least one sample product; the training sample set is constructed according to the user search data, and the specific implementation process can be as follows:

[0171] Obtaining sample search results corresponding to the sample search information, and determining a first sample product corresponding to at least one conversion factor in the sample search results, and a second sample product that does not have at least one conversion factor;

[0172] A training sample set is constructed according to the sample search information, the first sample commodity and the second sample commodity.

[0173] In an optional implementation of this embodiment, in addition to samples that have been exposed but have not performed operations related to conversion factors as negative samples, samples that have been highly exposed multiple times but have very low conversion rates can also be added as negative samples, that is, a training sample set is constructed based on sample search information, the first sample product, and the second sample product. The specific implementation process can be as follows:

[0174] Determine reference sample commodities whose commodity conversion rate is less than the conversion rate threshold value among the first sample commodities, and obtain reference sample attribute information of the reference sample commodities;

[0175] Acquire first sample attribute information of sample commodities other than the reference sample commodity in the first sample commodities, and second sample attribute information of the second sample commodities;

[0176] A positive sample is constructed according to the sample search information and the first sample attribute information, and a negative sample is constructed according to the sample search information and the second sample attribute information and the reference sample attribute information, respectively, to obtain a training sample set.

[0177] In an optional implementation of this embodiment, different samples may be resampled and weighted accordingly according to conversion factors such as clicks, inquiries, orders, and settlements. That is, after constructing a positive sample according to the sample search information and the first sample attribute information, the following may also be included:

[0178] Determine a sampling weight corresponding to a conversion factor of a target sample product, wherein the target sample product is any sample product in the first sample product except the reference sample product;

[0179] According to the sampling weights, the positive samples constructed based on the target sample products are resampled to obtain expanded positive samples.

[0180] Step 312: According to the correlation scores of each candidate data pair, at least one candidate product is divided into a first set numerical number of gears, the candidate products in each gear are sorted based on the set sorting rules, and at least one candidate product is displayed based on the sorting result and the gear order of the first set numerical number of gears.

[0181] Step 314: Divide at least one candidate product into a second set number of categories based on the category information of each candidate data pair; sort the candidate products in each category based on the relevance score; and display at least one candidate product based on the sorting result and the display order of the second set number of categories.

[0182] An embodiment of the present specification provides an object search method, which can obtain target search information and candidate attribute information of at least one candidate product; construct corresponding candidate data pairs based on the target search information and the candidate attribute information; input the candidate data pairs into a product analysis model to obtain a relevance score of the candidate data pairs, wherein the product analysis model is trained by a training sample set obtained based on at least one conversion factor sampling; and determine the display order of at least one candidate product and display it based on the relevance score of the candidate data pairs. In this case, a training sample set can be obtained in advance based on at least one conversion factor sampling, and a product analysis model can be obtained based on the training sample set. When the user inputs the target search information for search, the corresponding candidate data pairs can be constructed based on the target search information and the candidate attribute information of at least one candidate product. That is, the candidate data pair corresponding to a candidate product includes the target search information and the candidate attribute information of the candidate product. Each candidate data pair is input into the trained product analysis model, and the correlation score of each candidate data pair can be obtained. The correlation score can indicate the degree of correlation between the candidate attribute information and the target search information in the candidate data pair. Based on the correlation score of each candidate data pair, the display order of at least one candidate product can be determined and displayed to the user, thereby feeding back the search results to the user. In this way, at least one conversion factor can indicate the conversion rate of different dimensions, and the product analysis model incorporates the conversion rate information during training, that is, when the product analysis model learns the correlation between the search information and the product attribute information in the training sample set, it adds the conversion rate information of the product. Therefore, the trained product analysis model can comprehensively analyze the correlation between the target search information and the candidate attribute information in the input candidate data, as well as the conversion rate of the candidate products, balance the conversion rate and correlation, improve the accuracy of the search results, and ensure the user experience.

[0183] Corresponding to the above method embodiment, this specification also provides an object search device embodiment, Figure 4 FIG. 1 is a schematic diagram showing the structure of an object search device provided by an embodiment of the present specification. Figure 4 As shown, the device comprises:

[0184] An acquisition module 402 is configured to acquire target search information and candidate attribute information of at least one candidate object;

[0185] A construction module 404 is configured to construct corresponding candidate data pairs according to the target search information and the candidate attribute information;

[0186] An obtaining module 406 is configured to input the candidate data pair into an object analysis model to obtain a relevance score of the candidate data pair, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor;

[0187] The determination module 408 is configured to determine and display the display order of at least one candidate object according to the relevance score of the candidate data pair.

[0188] Optionally, the candidate attribute information of the candidate object includes object category information, industry information and attribute keywords; the construction module 404 is further configured to:

[0189] According to preset splicing rules, the target search information is spliced ​​with the object category information, industry information, and attribute keywords of the first candidate object to obtain a candidate data pair corresponding to the target search information and the first candidate object, wherein the first candidate object is any one of the at least one candidate object.

[0190] Optionally, the object analysis model includes a feature analysis layer and an output layer; the acquisition module 406 is further configured to:

[0191] Inputting the first candidate data pair into the feature analysis layer of the object analysis model to obtain the encoding features of each character in the first candidate data pair, wherein the first candidate data pair is a candidate data pair constructed based on the candidate attribute information of any candidate object;

[0192] The target coding feature of the first candidate data pair is determined by fusing the coding feature of the first character in the first candidate data pair, the average coding feature of each character, and the maximum coding feature;

[0193] The target coding features are input into the output layer of the object analysis model to obtain the relevance score of the first candidate data pair.

[0194] Optionally, the determination module 408 is further configured to:

[0195] According to the relevance scores of each candidate data pair, at least one candidate object is divided into a first set value of gears;

[0196] Sort the candidate objects in each gear based on the set sorting rules;

[0197] At least one candidate object is displayed based on the sorting result and the gear sequence of the first set number of gears.

[0198] Optionally, the determination module 408 is further configured to:

[0199] According to the category information of each candidate data pair, at least one candidate object is divided into a second set value of categories;

[0200] Sort the candidates within each category based on their relevance scores;

[0201] At least one candidate object is displayed based on the sorting result and the display order of the second set value categories.

[0202] Optionally, the device further comprises a training module configured to:

[0203] Acquire user search data within a preset time, and construct a training sample set according to the user search data, wherein the training sample set includes positive samples and negative samples, and the positive samples and the negative samples are obtained by sampling based on at least one conversion factor;

[0204] Input the training sample set into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set;

[0205] Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated through the preset boundary loss function, the model parameters of the initial analysis model are adjusted according to the loss value, and the operation steps of obtaining user search data within the preset time period are returned until the training stop condition is reached, and the trained object analysis model is obtained.

[0206] Optionally, the user search data includes sample search information and corresponding sample search results, and the sample search results include at least one sample object; and the training module is further configured to:

[0207] Obtaining a sample search result corresponding to the sample search information, and determining a first sample object corresponding to at least one conversion factor in the sample search result, and a second sample object that does not have at least one conversion factor;

[0208] A training sample set is constructed according to the sample search information, the first sample object and the second sample object.

[0209] Optionally, the training module is further configured to:

[0210] Determine a reference sample object whose object conversion rate is less than a conversion rate threshold in the first sample object, and obtain reference sample attribute information of the reference sample object;

[0211] Acquire first sample attribute information of sample objects other than the reference sample object in the first sample objects, and second sample attribute information of the second sample objects;

[0212] A positive sample is constructed according to the sample search information and the first sample attribute information, and a negative sample is constructed according to the sample search information and the second sample attribute information and the reference sample attribute information, respectively, to obtain a training sample set.

[0213] Optionally, the training module is further configured to:

[0214] Determine a sampling weight corresponding to a conversion factor of a target sample object, wherein the target sample object is any sample object in the first sample object except the reference sample object;

[0215] According to the sampling weights, the positive samples constructed based on the target sample objects are resampled to obtain expanded positive samples.

[0216] Optionally, the acquisition module 402 is further configured to:

[0217] Get the object title of at least one object to be displayed;

[0218] At least one candidate object whose object title is related to the target search information is determined from the objects to be displayed.

[0219] An embodiment of the present specification provides an object search device, which can obtain target search information and candidate attribute information of at least one candidate object; construct corresponding candidate data pairs according to the target search information and the candidate attribute information; input the candidate data pairs into an object analysis model to obtain the relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained based on at least one conversion factor sampling; and determine the display order of at least one candidate object and display it according to the relevance score of the candidate data pairs. In this case, a training sample set can be obtained in advance based on at least one conversion factor sampling, and an object analysis model can be obtained by training based on the training sample set. When a user inputs target search information for searching, a corresponding candidate data pair can be constructed based on the target search information and the candidate attribute information of at least one candidate object, that is, the candidate data pair corresponding to a candidate object includes the target search information and the candidate attribute information of the candidate object. Each candidate data pair is input into the trained object analysis model, and the relevance score of each candidate data pair can be obtained. The relevance score can indicate the degree of relevance between the candidate attribute information and the target search information in the candidate data pair. Based on the relevance score of each candidate data pair, the display order of at least one candidate object can be determined and displayed to the user, thereby feeding back the search results to the user. In this way, at least one conversion factor can indicate the conversion rate of different dimensions, and the object analysis model incorporates the conversion rate information during training, that is, when the object analysis model learns the correlation between the search information and the object attribute information in the training sample set, it takes into account the conversion rate information of the object. Therefore, the trained object analysis model can comprehensively analyze the correlation between the target search information and the candidate attribute information in the input candidate data, as well as the conversion rate of the candidate object, balance the conversion rate and correlation, improve the accuracy of the search results, and ensure the user experience.

[0220] The above is a schematic scheme of an object search device of this embodiment. It should be noted that the technical scheme of the object search device and the technical scheme of the object search method described above are of the same concept, and the details not described in detail in the technical scheme of the object search device can be referred to the description of the technical scheme of the object search method described above.

[0221] Figure 5 The block diagram of a computing device according to an embodiment of the present specification is shown. The components of the computing device 500 include but are not limited to a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.

[0222] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 550 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0223] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 5 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0224] The computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 500 may also be a mobile or stationary server.

[0225] The processor 520 is used to execute the following computer executable instructions to implement the following method:

[0226] Acquire target search information and candidate attribute information of at least one candidate object;

[0227] According to the target search information and the candidate attribute information, a corresponding candidate data pair is constructed;

[0228] Inputting the candidate data pairs into an object analysis model to obtain a relevance score of the candidate data pairs, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor;

[0229] According to the relevance scores of the candidate data pairs, a display order of at least one candidate object is determined and displayed.

[0230] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the object search method described above are of the same concept, and the details not described in detail in the technical scheme of the computing device can be found in the description of the technical scheme of the object search method described above.

[0231] An embodiment of the present specification further provides a computer-readable storage medium storing computer instructions, which are used to implement the steps of any object search method when executed by a processor.

[0232] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the object search method described above are of the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the object search method described above.

[0233] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0234] Computer instructions include computer program codes, which may be in source code form, object code form, executable files, or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM), random access memories (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0235] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0236] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0237] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. An object search method, comprising: Acquire target search information and candidate attribute information of at least one candidate object; Constructing corresponding candidate data pairs according to the target search information and the candidate attribute information; The candidate data pair is input into an object analysis model to obtain a correlation score of the candidate data pair, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor; the training sample set includes: positive samples of sample objects with operations related to the conversion factor, and negative samples of sample objects without operations related to the conversion factor; the conversion factor represents various operation parameters of the object conversion rate of the candidate object, the object conversion rate refers to the frequency with which the candidate object is finally acquired by the user under the search condition of the target search information, the at least one conversion factor indicates the object conversion rate of different dimensions, the positive sample also includes an expanded positive sample, the expanded positive sample is obtained by resampling the positive sample based on a sampling weight, the sampling weight is determined according to the conversion depth corresponding to the conversion factor, and the conversion depth is obtained by sorting the user's click behavior in sequence according to the depth; the negative sample also includes sample objects with a low conversion rate; According to the relevance score of the candidate data pair, the display order of the at least one candidate object is determined and displayed.

2. The object search method according to claim 1, wherein the candidate attribute information of the candidate object includes object category information, industry information and attribute keywords; Constructing corresponding candidate data pairs according to the target search information and the candidate attribute information, including: According to preset splicing rules, the target search information is spliced ​​with the object category information, industry information, and attribute keywords of the first candidate object to obtain a candidate data pair corresponding to the target search information and the first candidate object, wherein the first candidate object is any one of the at least one candidate object.

3. The object search method according to claim 1, wherein the object analysis model comprises a feature analysis layer and an output layer; Inputting the candidate data pair into an object analysis model to obtain a relevance score of the candidate data pair includes: Inputting a first candidate data pair into a feature analysis layer of the object analysis model to obtain encoding features of each character in the first candidate data pair, wherein the first candidate data pair is a candidate data pair constructed based on candidate attribute information of any candidate object; The target coding feature of the first candidate data pair is determined by fusing the coding feature of the first character in the first candidate data pair, the average coding feature of each character, and the maximum coding feature; The target coding feature is input into the output layer of the object analysis model to obtain the relevance score of the first candidate data pair.

4. The object search method according to claim 1, wherein the display order of the at least one candidate object is determined and displayed according to the relevance score of the candidate data pair, comprising: According to the relevance score of each candidate data pair, the at least one candidate object is divided into a first set value of gears; Sort the candidate objects in each gear based on the set sorting rules; Based on the sorting result and the gear order of the first set number of gears, the at least one candidate object is displayed.

5. The object search method according to claim 1, determining the display order of the at least one candidate object and displaying it according to the relevance score of the candidate data pair, comprising: According to the category information of each candidate data pair, the at least one candidate object is divided into a second set value of categories; Sort the candidates within each category based on their relevance scores; Based on the sorting result and the display order of the second set number of categories, the at least one candidate object is displayed.

6. The object search method according to any one of claims 1 to 5, wherein the object analysis model is obtained by training using the following method: Acquire user search data within a preset time, and construct a training sample set based on the user search data; Inputting the training sample set into the initial analysis model to obtain the prediction score corresponding to each sample in the training sample set; Based on the prediction score and sample type of each sample, the loss value of the initial analysis model is calculated by a preset boundary loss function, the model parameters of the initial analysis model are adjusted according to the loss value, and the operation step of obtaining user search data within a preset time period is returned to execute until the training stop condition is reached, thereby obtaining a trained object analysis model.

7. The object search method according to claim 6, wherein the user search data comprises sample search information and corresponding sample search results, and the sample search results comprise at least one sample object; Constructing a training sample set according to the user search data includes: Obtaining a sample search result corresponding to the sample search information, and determining a first sample object corresponding to the at least one conversion factor in the sample search result, and a second sample object that does not have the at least one conversion factor; The training sample set is constructed according to the sample search information, the first sample object and the second sample object.

8. The object search method according to claim 7, constructing the training sample set according to the sample search information, the first sample object and the second sample object, comprising: Determine a reference sample object whose object conversion rate is less than a conversion rate threshold among the first sample objects, and obtain reference sample attribute information of the reference sample object; Acquire first sample attribute information of sample objects other than the reference sample object in the first sample objects, and second sample attribute information of the second sample object; The positive sample is constructed according to the sample search information and the first sample attribute information, and the negative sample is constructed according to the sample search information, the second sample attribute information, and the reference sample attribute information, respectively, to obtain the training sample set.

9. The object search method according to claim 8, after constructing the positive sample according to the sample search information and the first sample attribute information, further comprising: Determine a sampling weight corresponding to a conversion factor of a target sample object, wherein the target sample object is any sample object in the first sample objects except the reference sample object; According to the sampling weights, the positive samples constructed based on the target sample object are resampled to obtain expanded positive samples.

10. The object search method according to any one of claims 1 to 5, after acquiring the target search information and before acquiring the candidate attribute information of at least one candidate object, further comprising: Get the object title of at least one object to be displayed; At least one candidate object whose object title is related to the target search information is determined from the objects to be displayed.

11. An object search device, comprising: An acquisition module, configured to acquire target search information and candidate attribute information of at least one candidate object; A construction module, configured to construct corresponding candidate data pairs according to the target search information and the candidate attribute information; an acquisition module, configured to input the candidate data pair into an object analysis model to obtain a relevance score of the candidate data pair, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor; the training sample set includes: positive samples of sample objects with operations related to the conversion factor, and negative samples of sample objects without operations related to the conversion factor; the conversion factor represents various operation parameters of the object conversion rate of the candidate object, the object conversion rate refers to the frequency with which the candidate object is finally acquired by the user under the search condition of the target search information, the at least one conversion factor indicates the object conversion rate of different dimensions, the positive sample also includes an expanded positive sample, the expanded positive sample is obtained by resampling the positive sample based on the corresponding sampling weight to determine the sampling weight corresponding to the conversion factor of the sample object in the positive sample, the sampling weight is determined according to the conversion depth corresponding to the conversion factor, and the conversion depth is obtained by sorting the user's click behavior in sequence according to the depth; the negative sample also includes sample objects with a low conversion rate; The determination module is configured to determine and display the display order of the at least one candidate object according to the relevance score of the candidate data pair.

12. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the following method: Acquire target search information and candidate attribute information of at least one candidate object; Constructing corresponding candidate data pairs according to the target search information and the candidate attribute information; The candidate data pair is input into an object analysis model to obtain a correlation score of the candidate data pair, wherein the object analysis model is trained by a training sample set obtained by sampling based on at least one conversion factor; the training sample set includes: positive samples of sample objects with operations related to the conversion factor, and negative samples of sample objects without operations related to the conversion factor; the conversion factor represents various operation parameters of the object conversion rate of the candidate object, the object conversion rate refers to the frequency with which the candidate object is finally acquired by the user under the search condition of the target search information, the at least one conversion factor indicates the object conversion rate of different dimensions, the positive sample also includes an expanded positive sample, the expanded positive sample is obtained by resampling the positive sample based on the corresponding sampling weight to determine the sampling weight corresponding to the conversion factor of the sample object in the positive sample, the sampling weight is determined according to the conversion depth corresponding to the conversion factor, and the conversion depth is obtained by sorting the user's click behavior in sequence according to the depth; the negative sample also includes sample objects with a low conversion rate; According to the relevance score of the candidate data pair, the display order of the at least one candidate object is determined and displayed.

13. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the object search method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Search engine sorting method and system and search engine

    CN104021125A

  • Application retrieval method and apparatus, storage medium and terminal

    CN108255954A

  • A Chinese text sentiment analysis method based on deep learning

    CN109697232A