Search result sorting method, device, electronic device and storage medium

By constructing the correspondence between domain classification and text feature weight parameters in the Internet of Things search engine, the problem of difference in feature priority in different fields affecting the accuracy of search results sorting is solved, and more efficient search results sorting is achieved.

CN114020866BActive Publication Date: 2025-05-23SHANDONG KURUI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111296285.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-05-23
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the priority differences in characteristics in different fields in IoT search engines, resulting in the impact of the accuracy of search results sorting.

Method used

By pre-constructing the correspondence between different domain classifications and the weight parameters of each text feature, query the corresponding weight parameters according to the domain classification of the search results, calculate the semantic similarity between the search results and the query words, and determine the scores based on the similarity and sort them.

Benefits of technology

It improves the accuracy of search results sorting, can better combine the priority differences in characteristics in different fields, and improves the efficiency of users in searching information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114020866B_ABST
    Figure CN114020866B_ABST
Patent Text Reader

Abstract

The present application relates to a search result sorting method, device, electronic device and storage medium, which are applied to the field of search technology. The method includes: obtaining multiple search results according to a query word input by a user; according to the field classifications corresponding to the multiple search results, querying the correspondence between the preset different field classifications and the weight parameters of each text feature, and determining the weight parameters of each text feature corresponding to each search result; extracting feature data related to each text feature from each search result; for each search result, calculating the semantic similarity between each search result and the query word according to the feature data related to each text feature and the weight parameters of each text feature; determining the similarity between each search result and the query word according to the semantic similarity, and then determining the score of each search result according to the similarity for sorting. The present application can improve the accuracy of the sorting results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of search technology, and in particular to a search result sorting method, device, electronic device and storage medium. Background Art

[0002] In the search field, in order to facilitate users to quickly find the desired information, the obtained search results can be sorted.

[0003] In related technologies, search engine ranking schemes can generally be classified into two categories. One is to use an algorithm based on a single data feature for ranking, such as the Pagerank algorithm. The other is to use a machine learning method to train a model based on manually annotated data sets, and use the trained model to score the search results for ranking.

[0004] However, in the Internet of Things (IoT) search engine, user searches involve many features such as devices, scenarios, and fields. In different application contexts, the priorities of these features are different. The current sorting method uses relatively single features and does not consider the priority differences of features in different fields, which affects the accuracy of search result sorting. Therefore, how to sort search results based on the priority of features in different fields has become a problem that needs to be solved urgently. Summary of the invention

[0005] In order to solve the above technical problem or at least partially solve the above technical problem, the present application provides a search result sorting method, device, electronic device and storage medium.

[0006] According to a first aspect of the present application, a method for sorting search results is provided, comprising:

[0007] Obtain multiple search results based on the query words entered by the user;

[0008] According to the field classifications respectively corresponding to the multiple search results, query the correspondence between the preset different field classifications and the weight parameters of each text feature, and determine the weight parameters of each text feature corresponding to each search result;

[0009] Extracting feature data related to each text feature from each of the search results;

[0010] For each of the search results, calculating the semantic similarity between each of the search results and the query term according to the feature data related to each of the text features and the weight parameters of each of the text features;

[0011] Determining the similarity between each search result and the query word according to the semantic similarity between each search result and the query word;

[0012] Determining a score for each of the search results based on a similarity between each of the search results and the query term;

[0013] The multiple search results are sorted according to the score of each search result.

[0014] According to a second aspect of the present application, a search result sorting device is provided, comprising:

[0015] A search result acquisition module is used to acquire multiple search results according to the query words input by the user;

[0016] A parameter determination module, for querying the correspondence between the preset different field classifications and the weight parameters of each text feature according to the field classifications respectively corresponding to the multiple search results, and determining the weight parameters of each text feature corresponding to each search result;

[0017] A feature extraction module, used to extract feature data related to each text feature from each search result;

[0018] A semantic similarity calculation module, for calculating the semantic similarity between each search result and the query word according to feature data related to each text feature and weight parameters of each text feature for each search result;

[0019] A similarity determination module, used to determine the similarity between each search result and the query word according to the semantic similarity between each search result and the query word;

[0020] A score determination module, used to determine the score of each search result according to the similarity between each search result and the query word;

[0021] The sorting module is used to sort the multiple search results according to the score of each search result.

[0022] According to a third aspect of the present application, an electronic device is provided, comprising: a processor, the processor being configured to execute a computer program stored in a memory, the computer program implementing the search result sorting method described in the first aspect when executed by the processor.

[0023] According to a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for sorting search results described in the first aspect is implemented.

[0024] According to a fifth aspect of the present application, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to execute the search result ranking method described in the first aspect.

[0025] Compared with the prior art, the technical solution provided by the embodiments of the present application has the following advantages:

[0026] By pre-constructing the correspondence between different domain classifications and weight parameters of each text feature, when multiple search results are obtained, the correspondence is queried according to the domain classification corresponding to each search result to determine the weight parameters of each text feature corresponding to each search result, and then for each search result, the semantic similarity between each search result and the query word is calculated according to the feature data related to each text feature extracted from the search result and the weight parameters of each text feature, and the similarity between each search result and the query word is determined according to the semantic similarity between each search result and the query word, and then the score of each search result is determined according to the similarity between each search result and the query word, and the multiple search results are sorted according to the score of each search result. By adopting the above technical solution, the weight parameters matching each text feature are determined according to the domain classification of the search result to calculate the semantic similarity and then determine the similarity between the search result and the query word and the score of the search result, so that the difference in the priority of features in different fields is combined when sorting the search results, which is conducive to improving the accuracy of the sorting results. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0028] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0029] Figure 1 A flowchart of a method for sorting search results provided in an embodiment of the present application;

[0030] Figure 2 A flowchart of a method for sorting search results provided in another embodiment of the present application;

[0031] Figure 3 An example diagram showing the fusion of features at different levels to determine the search result score;

[0032] Figure 4A schematic diagram of the structure of a search result sorting device provided in an embodiment of the present application;

[0033] Figure 5 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0034] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0036] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0038] Figure 1 This is a flow chart of a search result sorting method provided in an embodiment of the present application. This method can be performed by a search result sorting device provided in an embodiment of the present application, wherein the device can be implemented using software and / or hardware, and can generally be integrated in electronic devices such as search engine servers and cloud servers. Figure 1 As shown, the search result sorting method may include the following steps:

[0039] Step 101: Obtain multiple search results according to a query word input by a user.

[0040] In an embodiment of the present disclosure, when a user inputs a query word through a search portal of a client (such as a negative one screen search portal of a smartphone, a search portal of a browser, etc.), a cross-source search can be performed based on the obtained search word to obtain multiple search results related to the query word from different applications (such as shopping apps, information apps, etc.).

[0041] Step 102, according to the field classifications respectively corresponding to the plurality of search results, query the correspondence between the preset different field classifications and the weight parameters of each text feature, and determine the weight parameters of each text feature corresponding to each search result.

[0042] Among them, text features may include but are not limited to titles (or names), tags, description information, function introduction information, summary content, keywords, etc.; field classification may include but are not limited to shopping, audio and video entertainment, news, encyclopedia, games, health, education, life, travel, medicine, etc.

[0043] In the disclosed embodiment, for each field classification, a correspondence between different text features and weight parameters can be pre-constructed, and the weight parameters corresponding to each text feature under different field classifications can be determined through experiments or experience. For example, for the medical field, the title and description information are set with larger weight parameters, and for the news field, the title, abstract content and keywords are set with higher weight parameters.

[0044] After obtaining the search results, the field classification corresponding to each search result can be determined first. For example, a field classification model can be pre-trained and used to determine the field classification corresponding to each search result. Then, according to the field classification corresponding to each search result, the correspondence between the pre-constructed different field classifications and the weight parameters of each text feature is queried to determine the weight parameters of each text feature corresponding to each search result.

[0045] Exemplarily, assuming that in a pre-constructed correspondence, the correspondence between text features and weight parameters is recorded as "text features (weight parameters)", where, under the category of audio and video entertainment, the corresponding relationship is: title (a1), label (a2), description information (a3), keyword (a4); under the category of medicine, the corresponding relationship is: title (b1), label (b2), description information (b3), keyword (b4), and assuming that the field classification corresponding to a search result is audio and video entertainment, then the weight parameters of each text feature corresponding to the search result are title (a1), label (a2), description information (a3), keyword (a4).

[0046] Step 103: extracting feature data related to each text feature from each search result.

[0047] In the disclosed embodiment, for each search result, feature data related to each text feature can be extracted from the search result. For example, if the text feature includes a title, the related feature data includes the title of the search result; if the text feature includes a tag, the related feature data includes the tag information of the search result.

[0048] Exemplarily, a pre-trained feature extraction model may be used to extract feature data related to text features from search results, or other methods that can extract feature data corresponding to text features from search results may be used, which is not limited in this application.

[0049] Step 104 : For each of the search results, the semantic similarity between each of the search results and the query term is calculated based on the feature data related to each of the text features and the weight parameters of each of the text features.

[0050] As an example, a semantic similarity prediction model can be pre-trained and used to predict the feature data of each text feature corresponding to each search result and the semantic similarity between them and the query terms. A weighted sum is then taken based on the predicted semantic similarity and the weight parameters of each text feature to obtain the semantic similarity between each search result and the query terms.

[0051] As an example, for each search result, the first semantic similarity between the feature data related to the each text feature and the query term can be calculated; a weighted sum is performed based on the first semantic similarity and the weight parameters of the each text feature to obtain the semantic similarity between each search result and the query term.

[0052] Among them, when calculating the first semantic similarity between the feature data related to each text feature and the query word, if the feature data of a certain text feature is at least one word, the word vector of each word can be obtained, and the word vector of the query word can be obtained, and then the similarity between the word vector of each word and the word vector of the query word is calculated as the first semantic similarity between the feature data and the query word; if the feature data of a certain text feature is a sentence or a paragraph, keywords can be extracted from the feature data, such as using the term frequency-inverse document frequency index (Term Frequency-Inverse Document Frequency, TF-IDF) algorithm to extract keywords in the feature data, and obtain the word vector of each keyword and the word vector of the query word, and calculate the similarity between the word vector of each keyword and the word vector of the query word as the first semantic similarity of the feature data. Afterwards, according to the first semantic similarity between the feature data of each text feature and the query word, and the weight parameter of the corresponding text feature, weighted summation is performed to obtain the semantic similarity between each search result and the query word.

[0053] It should be noted that in the embodiments of the present application, a word vector model can be trained based on large-scale public texts (such as public encyclopedia corpus data and search results) to obtain word vectors for vocabulary, or a bag-of-words model can be constructed based on large-scale public texts to obtain word vectors for vocabulary. Other methods of obtaining word vectors for vocabulary that are not mentioned in this application can also be used to obtain word vectors, and this application does not impose any restrictions on this.

[0054] Step 105: Determine the similarity between each search result and the query word according to the semantic similarity between each search result and the query word.

[0055] In the embodiment of the present application, after obtaining the semantic similarity between each search result and the query word, the similarity between each search result and the query word can be determined according to the semantic similarity between each search result and the query word.

[0056] Exemplarily, the semantic similarity between the search result and the query word may be directly used as the similarity between the search result and the query word.

[0057] Exemplarily, the semantic similarity between the search result and the query word may be weighted and calculated with other similarity indicators to obtain the similarity between the search result and the query word. For example, the semantic similarity between the search result and the query word may be weighted and summed with a score determined based on the BM25 algorithm to obtain the similarity between the search result and the query word.

[0058] Step 106: Determine a score for each search result based on the similarity between each search result and the query term.

[0059] In this embodiment, after the similarity between each search result and the query term is determined, a score for each search result may be determined according to the similarity.

[0060] Exemplarily, the similarity between the search result and the query term may be directly used as the score of the search result.

[0061] For example, the similarity between the search results and the query words can be weighted and calculated with other evaluation indicators to obtain the score of the search results. For example, the similarity between the search results and the query words and the score corresponding to the download volume of the search results can be weighted and summed to obtain the score of the search results.

[0062] Step 107: sort the multiple search results according to the score of each search result.

[0063] In this embodiment, after obtaining the score of each search result, the obtained search results can be sorted according to the score of each search result. For example, the search results can be sorted in the order of small to large scores, or in the order of large to small scores. The search engine server can display the sorting results in the client page held by the user.

[0064] The search result sorting method of this embodiment pre-constructs the correspondence between different field classifications and weight parameters of each text feature. When multiple search results are obtained, the correspondence is queried according to the field classification corresponding to each search result to determine the weight parameters of each text feature corresponding to each search result. Then, for each search result, the semantic similarity between each search result and the query word is calculated according to the feature data related to each text feature extracted from the search result and the weight parameters of each text feature. Then, according to the semantic similarity between each search result and the query word, the similarity between each search result and the query word is determined. Then, according to the similarity between each search result and the query word, the score of each search result is determined, and multiple search results are sorted according to the score of each search result. By adopting the above technical solution, the weight parameters matching each text feature are determined according to the field classification of the search result to calculate the semantic similarity and then determine the similarity between the search result and the query word and the score of the search result. This realizes that the difference in the priority of features in different fields is combined when sorting the search results, which is conducive to improving the accuracy of the sorting results and solving the technical problem of poor sorting accuracy caused by using a single feature for sorting in the prior art.

[0065] Generally, the download volume and score of the search results can also reflect the quality of the search results to a certain extent. The embodiment of the present application also proposes that the download volume and score of the search results can be used as reference factors for sorting. Therefore, in an optional implementation of the present application, the method further includes:

[0066] Obtain the download volume and rating of each of the search results;

[0067] Determining the popularity of each of the search results according to the download volume and the rating;

[0068] Accordingly, determining the score of each search result according to the similarity between each search result and the query term includes:

[0069] The score of each search result is determined according to the similarity between each search result and the query term, and the popularity of each search result.

[0070] In the embodiment of the present application, the download volume and rating of each search result can be obtained to evaluate the popularity of each search result. For example, when the search result is an application, the download volume and rating of the application can be obtained from the source of the application. When the search result is a web document, the download volume and rating of the web document can also be obtained from the source of the web document. When the search result is a web article that does not support downloading and evaluation, the number of views of the web article can be used as its download volume, and its rating is 0. The popularity of each search result can be determined using the obtained download volume and rating.

[0071] For example, the popularity of the search results can be calculated using the following formula (1).

[0072] Pop(item)=sigmoid(Ratings)+tanh(Downloads) (1)

[0073] Among them, Pop(item) indicates the popularity of the search results, Ratings indicates the rating of the search results, and Downloads indicates the download volume of the search results.

[0074] Afterwards, for each search result, a score of the search result may be determined based on the similarity between the search result and the query term, and the popularity of the search result.

[0075] Exemplarily, a weighted sum may be performed on the similarity between the search result and the query term and the popularity of the search result to obtain a score for the search result.

[0076] Exemplarily, the similarity between the search result and the query term and the popularity of the search result may be multiplied to obtain the score of the search result.

[0077] In an embodiment of the present application, the popularity of each search result is determined by obtaining the download volume and score of each search result, and then the score of each search result is determined based on the similarity between each search result and the query term, as well as the popularity of each search result. When calculating the score of the search result, not only the similarity between the search result and the query term is taken into account, but also the popularity of the search result is taken into account, which is conducive to improving the accuracy of the search result scoring, and then improving the accuracy of the search result sorting.

[0078] Furthermore, in an optional implementation manner of the present application, the method further includes:

[0079] Determine the popularity of each application to which the search result belongs according to the first ranking data of the application to which the search result belongs;

[0080] Determining the intent similarity between the query term and the application to which each search result belongs according to the intent corresponding to the application and the intent corresponding to the query term;

[0081] Extracting at least one application keyword from the text information of the application, and determining a second semantic similarity between the at least one application keyword and the query word according to a word vector similarity between the at least one application keyword and the query word and an inverse text frequency index of the at least one application keyword;

[0082] Determining a first relevance score between the query term and the text information of the application program using the BM25 algorithm;

[0083] Determining the similarity between the application and the query term according to the first relevance score and the second semantic similarity;

[0084] Determine a score of the application to which each of the search results belongs according to the popularity of the application, the intention similarity between the query word and the application to which each of the search results belongs, and the similarity between the application to which each of the search results belongs and the query word;

[0085] Determining the score of each search result according to the similarity between each search result and the query term and the popularity of each search result includes:

[0086] The score of each search result is determined according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

[0087] It is understandable that no matter the search results are online articles, shopping websites, music, or recommended APP download addresses, or other types of content, they all have corresponding sources, which are from a certain application or a certain channel of a certain application. For example, the recommended APP comes from the application store, the online article comes from the entertainment news channel of the browser, etc. The application from which the search results come is the application to which the search results belong, and the channel from which the search results come is called the channel to which the search results belong.

[0088] In the embodiment of the present application, for each search result, ranking data (called first ranking data) of the application to which each search result belongs can be obtained, and the popularity of the application to which each search result belongs can be determined based on the first ranking data.

[0089] For example, the ranking data of the application can be obtained from the network, and the ranking data can be normalized into a value ranging from 1 to 100 to represent the popularity of the application. In addition, in some embodiments, the download volume, rating and other information of the application can also be obtained, and combined with the ranking data of the application to normalize a specific value as the popularity of the application, which is not limited in this application.

[0090] In the embodiment of the present application, the intent corresponding to the query word and the intent corresponding to the application to which the search result belongs can also be analyzed. For example, the intent corresponding to the query word can be identified by a pre-trained intent recognition model, and the intent satisfied by the application to which the search result belongs can be determined based on the name and description information of the application to which the search result belongs. Afterwards, based on the intent corresponding to the application to which the search result belongs and the intent corresponding to the query word, the intent similarity between the query word and the application to which each search result belongs can be determined.

[0091] Exemplarily, the intent similarity between the query term and the application to which each search result belongs can be calculated using the following formula (2).

[0092] P(APP intent |query intent )=log(1+β 1 +cos(APP intent |query intent )) (2)

[0093] Among them, P(APP intent |query intent ) represents the intent similarity between the query term and the application to which each search result belongs, β 1 is a hyperparameter whose value can be preset, cos(APP intent |query intent) represents the cosine similarity between the word vector of the intent corresponding to the query word and the word vector of the intent corresponding to the application to which it belongs.

[0094] In the embodiment of the present application, the text information of the application to which it belongs can be obtained, and the text information may include but is not limited to the name, introduction information, function description, label and other information. Then, a commonly used keyword extraction algorithm can be used to extract at least one keyword from the text information of the application to which it belongs as at least one application keyword. After that, the word vector of at least one application keyword is obtained, and the word vector similarity between the word vector of each application keyword and the word vector of the query word is calculated, and the inverse document frequency index (IDF) value of each application keyword is used as the weight value to perform weighted summation calculation on the word vector similarity between the word vector of each application keyword and the word vector of the query word, and obtain the second semantic similarity between each application keyword and the query word. Then, the BM25 algorithm is used to determine the first correlation score between the query word and the text information of the application to which it belongs, wherein the BM25 algorithm is a commonly used algorithm for determining the correlation between the query word and the document, and the specific processing process is not described in detail in this application. In the embodiment of the present application, according to the first correlation score and the second semantic similarity, the similarity between the application to which each search result belongs and the query word can be determined.

[0095] For example, the similarity between the application to which each search result belongs and the query term can be calculated using the following formula (3).

[0096] P(APP|query)=log(1+Sim(APP,query)) (3)

[0097] Among them, P(APP|query) represents the similarity between the application and the query, and Sim(APP,query) is determined by the first relevance score and the second semantic similarity. For example, for each search result, the first relevance score between the query determined by the BM25 algorithm and the text information of the application to which the search result belongs, and the second semantic similarity between the application keywords related to the application and the query can be weighted and summed to obtain Sim(APP,query), and the weighting coefficient can be preset. Of course, other methods such as averaging can also be used to determine Sim(APP,query), and this application does not limit this.

[0098] Furthermore, for each search result, the score of the application to which each search result belongs can be determined based on the popularity of the application to which the search result belongs, the intent similarity between the query term and the application to which the search result belongs, and the similarity between the application and the query term.

[0099] Exemplarily, the score of the application to which each search result belongs can be calculated by the following formula (4).

[0100] Score(APP,query)=Pop(APP)*P(APP|query)*P(APP intent |query intent ) (4)

[0101] Where Score(APP,query) represents the score of the application to which the search result belongs, Pop(APP) represents the popularity of the application to which the search result belongs, P(APP|query) is the similarity between the application and the query term calculated by the above formula (3), and P(APP intent |query intent ) is the intent similarity between the query term and the application to which the search result belongs, calculated by the above formula (2).

[0102] Furthermore, determining the score of each search result according to the similarity between each search result and the query term and the popularity of each search result includes:

[0103] The score of each search result is determined according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

[0104] In an embodiment of the present application, when determining the score of each search result, the score of the search result can be determined based on the similarity between the search result and the query term and the popularity of the search result, combined with the score of the application to which the search result belongs.

[0105] Exemplarily, the similarity between the search result and the query term and the popularity of the search result may be multiplied and then added to the score of the application to which the search result belongs to obtain the score of the search result.

[0106] In an embodiment of the present application, the score of each search result is determined based on the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result. This ensures that when calculating the score of the search result, not only the characteristics of the search result itself but also the characteristics of the application to which it belongs are taken into account, and the mutual influence of different levels is fully considered, which is conducive to improving the accuracy of the sorting.

[0107] Furthermore, in an optional implementation manner of the present application, the method further includes:

[0108] Determine the popularity of the channel according to the first ranking data of the application to which each search result belongs and the second ranking data of the channel in the application;

[0109] Determine the intent similarity between the query word and the channel to which each search result belongs according to the intent corresponding to the channel and the intent corresponding to the query word;

[0110] Extracting at least one channel keyword from the text information of the channel, and determining a third semantic similarity between the at least one channel keyword and the query word;

[0111] Determine a second correlation score between the query word and the text information of the channel using the BM25 algorithm;

[0112] Determining the similarity between the belonging channel in the belonging application and the query word according to the second relevance score, the third semantic similarity, and the similarity between the belonging application and the query word;

[0113] Determine a score of the channel to which each search result belongs according to the popularity of the channel, the intention similarity between the query word and the channel to which each search result belongs, and the similarity between the channel and the query word;

[0114] Determining the score of each search result according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result includes:

[0115] The score of each search result is determined according to the score of the channel to which each search result belongs, the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

[0116] In an embodiment of the present application, for each search result, the ranking data of the application to which each search result belongs (referred to as the first ranking data) and the ranking data of the channel to which each search result belongs in the application to which it belongs (referred to as the second ranking data) can be obtained, and the popularity of the channel to which the search result belongs can be determined based on the first ranking data and the second ranking data.

[0117] For example, the first ranking data of the application and the second ranking data of the channel can be obtained from the network, and the first ranking data and the second ranking data can be normalized into a value ranging from 1 to 100 to represent the popularity of the channel. In addition, in some embodiments, the download volume, rating and other information of the application to which the channel belongs can also be obtained, and the ranking data of the application and the ranking data of the channel can be combined to normalize and obtain a specific value as the popularity of the channel, which is not limited in this application.

[0118] In the embodiment of the present application, the intent corresponding to the query word and the intent corresponding to the channel to which the search results belong may also be analyzed. For example, the intent corresponding to the query word may be identified by a pre-trained intent recognition model, and the intent satisfied by the channel may be determined based on the name and description information of the channel. Thereafter, based on the intent corresponding to the channel and the intent corresponding to the query word, the intent similarity between the query word and the channel to which each search result belongs may be determined.

[0119] Exemplarily, the intent similarity between the query term and the channel to which each search result belongs can be calculated using the following formula (5).

[0120] P(Channel intent |q intent )=log(1+β+cos(Channel intent |query intent )) (5)

[0121] Among them, P (Channel intent |q intent ) represents the intent similarity between the query term and the channel to which each search result belongs, β 2 is a hyperparameter whose value can be preset. intent |query intent ) represents the cosine similarity between the word vector of the intent corresponding to the query word and the word vector of the intent corresponding to the channel to which it belongs.

[0122] In the embodiment of the present application, the text information of the channel can be obtained, and the text information may include but is not limited to the name, introduction information, function description, label and other information. Then, a commonly used keyword extraction algorithm can be used to extract at least one keyword from the text information of the channel as at least one channel keyword. According to the word vector of at least one channel keyword and the word vector of the query word, the third semantic similarity between the at least one channel keyword and the query word can be calculated, wherein the third semantic similarity is obtained by weighted summing the word vector similarity between the word vector of each channel keyword and the word vector of the query word, and the IDF value of each channel keyword is the weight value when the weighted sum is calculated. Then, the BM25 algorithm is used to determine the second correlation score between the query word and the text information of the channel to which it belongs. In the embodiment of the present application, according to the second correlation score, the third semantic similarity and the similarity between the application to which it belongs and the query word, the similarity between the channel to which it belongs and the query word can be determined. Among them, the similarity between the application to which it belongs and the query word has been calculated in the aforementioned embodiment.

[0123] Exemplarily, for each search result, a weighted sum calculation may be performed on the relevant second relevance score, the third semantic similarity, and the similarity between the corresponding application and the query term to obtain the similarity between the corresponding channel in the corresponding application and the query term.

[0124] Furthermore, for each search result, the score of the channel to which each search result belongs can be determined based on the popularity of the channel to which the search result belongs, the intent similarity between the query term and the channel to which the search result belongs, and the similarity between the channel and the query term.

[0125] Exemplarily, the score of the channel to which each search result belongs can be calculated by the following formula (6).

[0126]

[0127] Among them, Score(Channel,query) represents the score of the channel to which the search result belongs, Pop(Channel) represents the popularity of the channel to which the search result belongs, P(Channel|query) represents the similarity between the channel and the query term, and P(Channel intent |q intent ) is the intent similarity between the query term and the channel to which the search result belongs, calculated by the above formula (5).

[0128] Furthermore, when determining the score of any search result, the score of the search result may be determined based on the score of the channel to which the search result belongs, the score of the application to which the search result belongs, the similarity between the search result and the query term, and the popularity of the search result.

[0129] Exemplarily, the score of the search result may be calculated using the following formula (7).

[0130]

[0131] Among them, Score(item, query) represents the score of the search result, Pop(item) represents the popularity of the search result, P(item|query) represents the similarity between the search result and the query term, Score(Channel, query) represents the score of the channel to which the search result belongs, Score(APP, query) represents the score of the application to which the search result belongs, α, β and γ are hyperparameters, and the values ​​can be set in advance.

[0132] It can be understood that each part of the formula involved in the embodiments of the present application can be set with corresponding parameters, and the values ​​of the parameters can be determined in advance through experiments.

[0133] In an embodiment of the present application, when determining the score of a search result, the score of each search result is determined based on the score of the channel to which each search result belongs, the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result. This achieves the integration of features at different levels (search results, channels to which they belong, applications to which they belong) to determine the score of the search result, fully considers the mutual influence of different levels, and is conducive to improving the accuracy of search result sorting.

[0134] In an optional implementation of the present application, when determining the similarity between the search results and the query, the similarity between the application to which the search results belong and the query, and the similarity between the channel to which the search results belong and the query may be combined to determine the similarity, so as to improve the accuracy. Figure 2 As shown, in Figure 1 Based on the illustrated embodiment, step 105 may include the following sub-steps:

[0135] Step 201: Determine a third relevance score between each of the search results and the query term using the BM25 algorithm.

[0136] It can be understood that the BM25 algorithm is a commonly used algorithm for determining the relevance between a query term and a document, so in this embodiment, for each search result, the BM25 algorithm can be used to determine the relevance between each search result and the query term to obtain a third relevance score. One search result corresponds to one third relevance score.

[0137] Step 202, determining the similarity between each search result and the query word according to the semantic similarity between each search result and the query word, the similarity between the belonging channel and the query word, the similarity between the belonging application and the query word, and the third relevance score.

[0138] Among them, the calculation methods of the semantic similarity between the search results and the query words, the similarity between the corresponding channels and the query words, and the similarity between the corresponding applications and the query words have been described in detail in the above embodiments and will not be repeated here.

[0139] For example, for each search result, a weighted sum calculation may be performed on the calculated semantic similarity between the search result and the query, the similarity between the channel to which the search result belongs and the query, the similarity between the application to which the search result belongs and the query, and the third relevance score determined by the BM25 algorithm to score the similarity between the search result and the query. The weight value of each part may be preset.

[0140] The search results of this embodiment are sorted by using the BM25 algorithm to determine the third correlation score between each search result and the query term, and then determining the similarity between each search result and the query term based on the semantic similarity between the search result and the query term, the similarity between the channel to which the search result belongs and the query term, the similarity between the application to which the search result belongs and the query term, and the third correlation score. As a result, when calculating the similarity between the search results and the query term, not only the correlation between the search results themselves and the query term is considered, but also the similarity between the application to which the search result belongs and the channel to which the search term is integrated. By integrating multiple features to calculate the similarity between the search results and the query term, it is beneficial to improve the scoring accuracy of the search results, and thus improve the accuracy of the sorting of the search results.

[0141] In the embodiment of the present application, the application, the frequency band of the application and the features of the search result content are integrated with each other. Figure 3 An example diagram showing how to combine features at different levels to determine the search result score is shown in FIG. Figure 3As shown, the score of the application is determined based on the structured features of the application (such as the popularity of the application, the intents satisfied by the application, the title, tags, etc.) and (unstructured features); the score of the channel is determined based on the structured features of the channel (such as the popularity of the channel, the intents satisfied by the channel, the title, tags, etc.) and semantic similarity (unstructured features), wherein the score of the application is integrated into the structured features of the channel; the score of the content is determined based on the structured features of the search result content (such as the popularity of the content, the intents satisfied by the content, the title, tags, etc.) and semantic similarity (unstructured features), wherein the score of the application and the score of the channel are integrated into the structured features of the search result content. It can be seen that the solution provided by the present application integrates the relationship features between different levels (search results, channels, applications), fully considers the mutual influence of different levels, and the features of different levels have corresponding parameters in the above calculation formula. The value of the parameter of each feature can be set according to the feature priority, which not only ensures that the content in different fields can be uniformly calculated through an integrated ranking formula, but also facilitates the intuitive addition, deletion and adjustment of the importance of features in different scenarios (fields).

[0142] Corresponding to the above method embodiment, the embodiment of the present application also provides a search result sorting device.

[0143] Figure 4 A schematic diagram of the structure of a search result sorting device provided in an embodiment of the present application, such as Figure 4 As shown, the search result sorting device 30 may include: a search result acquisition module 310, a parameter determination module 320, a feature extraction module 330, a semantic similarity calculation module 340, a similarity determination module 350, a score determination module 360 ​​and a sorting module 370.

[0144] The search result acquisition module 310 is used to acquire multiple search results according to the query word input by the user;

[0145] The parameter determination module 320 is used to query the correspondence between the preset different field classifications and the weight parameters of each text feature according to the field classifications corresponding to the multiple search results, and determine the weight parameters of each text feature corresponding to each search result;

[0146] A feature extraction module 330, used to extract feature data related to each text feature from each search result;

[0147] A semantic similarity calculation module 340, configured to calculate, for each of the search results, a semantic similarity between each of the search results and the query term based on feature data related to each of the text features and weight parameters of each of the text features;

[0148] A similarity determination module 350, configured to determine the similarity between each search result and the query term according to the semantic similarity between each search result and the query term;

[0149] A score determination module 360, configured to determine a score for each of the search results according to a similarity between each of the search results and the query term;

[0150] The sorting module 370 is used to sort the multiple search results according to the score of each search result.

[0151] Optionally, the semantic similarity calculation module 340 is specifically used to:

[0152] For each of the search results, calculating a first semantic similarity between the feature data related to each of the text features and the query term;

[0153] A weighted sum is performed according to the first semantic similarity and the weight parameters of each text feature to obtain the semantic similarity between each search result and the query word.

[0154] Optionally, the search result sorting device 30 further includes:

[0155] A rating acquisition module, used to obtain the download volume and rating of each search result;

[0156] A search result popularity determination module, used to determine the popularity of each search result according to the download volume and the rating;

[0157] Accordingly, the score determination module 360 ​​is further configured to:

[0158] The score of each search result is determined according to the similarity between each search result and the query term, and the popularity of each search result.

[0159] Optionally, the search result sorting device 30 further includes:

[0160] An application popularity determination module, used to determine the popularity of the application to which each search result belongs according to the first ranking data of the application to which the search result belongs;

[0161] A first intent similarity determination module, configured to determine the intent similarity between the query term and the application to which each search result belongs, based on the intent corresponding to the application and the intent corresponding to the query term;

[0162] An application semantic similarity determination module is used to extract at least one application keyword from the text information of the application program to which it belongs, and

[0163] Determining a second semantic similarity between the at least one application keyword and the query word according to the word vector similarity between the at least one application keyword and the query word and the inverse text frequency index of the at least one application keyword;

[0164] A first relevance determination module, configured to determine a first relevance score between the query term and the text information of the application program by using a BM25 algorithm;

[0165] an application similarity determination module, configured to determine the similarity between the application and the query term according to the first relevance score and the second semantic similarity;

[0166] An application score determination module, configured to determine a score of the application to which each search result belongs according to the popularity of the application to which it belongs, the intention similarity between the query term and the application to which each search result belongs, and the similarity between the application to which it belongs and the query term;

[0167] Accordingly, the score determination module 360 ​​is further configured to:

[0168] The score of each search result is determined according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

[0169] Optionally, the search result sorting device 30 further includes:

[0170] A channel popularity determination module, configured to determine the popularity of the channel according to the first ranking data of the application to which each search result belongs and the second ranking data of the channel to which each search result belongs in the application;

[0171] A second intent similarity determination module, configured to determine the intent similarity between the query term and the channel to which each search result belongs according to the intent corresponding to the channel and the intent corresponding to the query term;

[0172] A channel semantic similarity determination module, configured to extract at least one channel keyword from the text information of the channel, and determine a third semantic similarity between the at least one channel keyword and the query word;

[0173] A second relevance determination module, configured to determine a second relevance score between the query word and the text information of the channel to which the query word belongs using the BM25 algorithm;

[0174] a channel similarity determination module, configured to determine the similarity between the channel in the application and the query word according to the second relevance score, the third semantic similarity, and the similarity between the application and the query word;

[0175] A channel score determination module, configured to determine a score of the channel to which each search result belongs according to the popularity of the channel, the intention similarity between the query word and the channel to which each search result belongs, and the similarity between the channel and the query word;

[0176] Accordingly, the score determination module 360 ​​is further configured to:

[0177] The score of each search result is determined according to the score of the channel to which each search result belongs, the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

[0178] Optionally, the similarity determination module 350 includes:

[0179] A third relevance determination unit, configured to determine a third relevance score between each of the search results and the query term using a BM25 algorithm;

[0180] A similarity determination unit is used to determine the similarity between each search result and the query word according to the semantic similarity between each search result and the query word, the similarity between the belonging channel and the query word, the similarity between the belonging application and the query word, and the third relevance score.

[0181] The search result sorting device provided in the embodiments of the present disclosure can execute any search result sorting method provided in the embodiments of the present disclosure that can be applied to electronic devices such as search engine servers, and has functional modules and beneficial effects corresponding to the execution method. Contents not fully described in the embodiments of the present disclosure device can refer to the description in any method embodiment of the present disclosure.

[0182] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0183] In an exemplary embodiment of the present application, an electronic device is further provided, comprising: a processor, wherein the processor is configured to execute a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the steps of the search result sorting method as described in the above embodiment.

[0184] Figure 5 This is a schematic diagram of a structure of an electronic device provided in an embodiment of the present application. It should be noted that: Figure 5 The electronic device 500 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0185] like Figure 5 As shown, electronic device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In RAM 503, various programs and data required for system operation are also stored. Central processing unit 501, ROM 502 and RAM 503 are connected to each other via a bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0186] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage section 508 as needed.

[0187] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 509, and / or installed from a removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the device of the present application are executed.

[0188] In an embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the search result sorting method described in the above embodiment are implemented.

[0189] It should be noted that the computer-readable storage medium shown in the present application may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, radio frequency, etc., or any suitable combination of the above.

[0190] In an embodiment of the present application, a computer program product is further provided. When the computer program product is executed on a computer, the computer is enabled to execute the steps of the search result sorting method described in the above embodiment.

[0191] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0192] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for sorting search results, It is characterized in that The method comprises: Obtain multiple search results based on the query words entered by the user; According to the field classifications respectively corresponding to the multiple search results, query the correspondence between the preset different field classifications and the weight parameters of each text feature, and determine the weight parameters of each text feature corresponding to each search result; Extracting feature data related to each text feature from each of the search results; For each of the search results, calculating the semantic similarity between each of the search results and the query term according to the feature data related to each of the text features and the weight parameters of each of the text features; Determining the similarity between each search result and the query word according to the semantic similarity between each search result and the query word; Determining a score for each of the search results based on a similarity between each of the search results and the query term; Sorting the multiple search results according to the score of each search result; The step of calculating the semantic similarity between each search result and the query term according to the feature data related to each text feature and the weight parameter of each text feature for each search result includes: For each of the search results, calculating a first semantic similarity between the feature data related to each of the text features and the query term; A weighted sum is performed according to the first semantic similarity and the weight parameters of each text feature to obtain the semantic similarity between each search result and the query word.

2. The method for ranking search results according to claim 1, It is characterized in that The method further comprises: Obtain the download volume and rating of each of the search results; Determining the popularity of each of the search results according to the download volume and the rating; Accordingly, determining the score of each search result according to the similarity between each search result and the query term includes: The score of each search result is determined according to the similarity between each search result and the query term, and the popularity of each search result.

3. The method for ranking search results according to claim 2, It is characterized in that The method further comprises: Determine the popularity of each application to which the search result belongs according to the first ranking data of the application to which the search result belongs; Determining the intent similarity between the query term and the application to which each search result belongs according to the intent corresponding to the application and the intent corresponding to the query term; Extracting at least one application keyword from the text information of the application, and determining a second semantic similarity between the at least one application keyword and the query word according to a word vector similarity between the at least one application keyword and the query word and an inverse text frequency index of the at least one application keyword; Determine a first relevance score between the query term and the text information of the application program using a BM25 algorithm; Determining the similarity between the application and the query term according to the first relevance score and the second semantic similarity; Determine a score of the application to which each of the search results belongs according to the popularity of the application, the intention similarity between the query word and the application to which each of the search results belongs, and the similarity between the application to which each of the search results belongs and the query word; Determining the score of each search result according to the similarity between each search result and the query term and the popularity of each search result includes: The score of each search result is determined according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

4. The method for ranking search results according to claim 3, It is characterized in that The method further comprises: Determine the popularity of the channel according to the first ranking data of the application to which each of the search results belongs and the second ranking data of the channel to which each of the search results belongs in the application to which it belongs; Determine the intent similarity between the query word and the channel to which each search result belongs according to the intent corresponding to the channel and the intent corresponding to the query word; Extracting at least one channel keyword from the text information of the channel, and determining a third semantic similarity between the at least one channel keyword and the query word; Determine a second correlation score between the query word and the text information of the channel using the BM25 algorithm; Determining the similarity between the belonging channel in the belonging application and the query word according to the second relevance score, the third semantic similarity, and the similarity between the belonging application and the query word; Determine a score of the channel to which each search result belongs according to the popularity of the channel, the intention similarity between the query word and the channel to which each search result belongs, and the similarity between the channel and the query word; Determining the score of each search result according to the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result includes: The score of each search result is determined according to the score of the channel to which each search result belongs, the score of the application to which each search result belongs, the similarity between each search result and the query term, and the popularity of each search result.

5. The method for ranking search results according to claim 4, It is characterized in that Determining the similarity between each search result and the query word according to the semantic similarity between each search result and the query word includes: Determine a third relevance score between each of the search results and the query term using the BM25 algorithm; The similarity between each search result and the query word is determined according to the semantic similarity between each search result and the query word, the similarity between the belonging channel and the query word, the similarity between the belonging application and the query word, and the third relevance score.

6. A search result sorting device, It is characterized in that include: A search result acquisition module is used to acquire multiple search results according to the query words input by the user; A parameter determination module, for querying the correspondence between the preset different field classifications and the weight parameters of each text feature according to the field classifications respectively corresponding to the multiple search results, and determining the weight parameters of each text feature corresponding to each search result; A feature extraction module, used to extract feature data related to each text feature from each search result; A semantic similarity calculation module, for calculating the semantic similarity between each search result and the query word according to feature data related to each text feature and weight parameters of each text feature for each search result; A similarity determination module, used to determine the similarity between each search result and the query word according to the semantic similarity between each search result and the query word; A score determination module, used to determine the score of each search result according to the similarity between each search result and the query word; A sorting module, used to sort the multiple search results according to the score of each search result; Wherein, the semantic similarity calculation module is specifically used for: For each of the search results, calculating a first semantic similarity between the feature data related to each of the text features and the query term; A weighted sum is performed according to the first semantic similarity and the weight parameters of each text feature to obtain the semantic similarity between each search result and the query word.

7. An electronic device, It is characterized in that include: A processor, wherein the processor is used to execute a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the steps of the search result sorting method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the search result ranking method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Search result providing method and device

    CN104008170A

  • Data processing method and device, terminal device, and computer storage medium

    CN109376298A