Recall result ranking method and apparatus, storage medium, device, and program product

CN122838702APending Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373636.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,相关技术中的召回结果排序方法主要依赖人工标注或基于规则的评分机制,准确性方面仍存在诸多不足

Benefits of technology

[0052]本申请实施例提供的获取查询信息、与查询信息对应的查询对象的对象画像信息,以及与查询信息对应的多个召回结果;基于查询信息与对象画像信息,对多个召回结果中的各个召回结果进行特征标注,得到各个召回结果的标注信息,标注信息用于指示各个召回结果与查询信息和/或对象画像信息的匹配程度;基于各个召回结果的标注信息、查询信息与对象画像信息,对多个召回结果进行排序的技术方案,通过综合考虑查询信息、标注信息、对象画像信息对召回结果进行排序处理,更精准地了解查询对象的特征、偏好、历史行为等,提高召回结果排序的准确性,此外,标注信息明确地指示了召回结果与查询信息和/或对象画像信息的匹配程度,提高了召回结果的排序结果的可解释性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838702A_ABST
    Figure CN122838702A_ABST
Patent Text Reader

Abstract

The application discloses a recall result sorting method and device, a storage medium, equipment and a program product, and is applied to scenes such as a search engine, news recommendation, artificial intelligence and e-commerce recommendation. The method comprises the following steps: acquiring query information, object portrait information of a query object corresponding to the query information, and a plurality of recall results corresponding to the query information; based on the query information and the object portrait information, the plurality of recall results are marked with features, and the marked information of each recall result is obtained, the marked information is used for indicating the matching degree of each recall result and the query information and / or the object portrait information; and based on the marked information of each recall result, the query information and the object portrait information, the plurality of recall results are sorted. The application sorts the plurality of recall results in combination with the object portrait information of the query object and the query information, and improves the accuracy and personalization of the recall result sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, specifically to a method, apparatus, storage medium, device, and program product for sorting recall results. Background Technology

[0002] With the rapid development of internet technology, information retrieval and recommendation systems are playing an increasingly important role in modern society. Users access massive amounts of information daily through platforms such as search engines, news apps, and social media. Filtering the most relevant and valuable content from this information has become one of the core challenges of information retrieval and recommendation systems. Result ranking, as a key technology in information retrieval and recommendation systems, directly impacts the user's search experience and information acquisition efficiency.

[0003] However, the recall result ranking methods in related technologies mainly rely on manual annotation or rule-based scoring mechanisms, which still have many shortcomings in terms of accuracy. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, device, and program product for ranking recall results, which can accurately understand the characteristics, preferences, and historical behavior of the query object, thereby improving the accuracy of recall result ranking. In addition, the annotation information clearly indicates the degree of matching between the recall result and the query information and / or object profile information, thereby improving the interpretability of the recall result ranking.

[0005] On one hand, embodiments of this application provide a method for ranking recall results, the method comprising:

[0006] Obtain query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information;

[0007] Based on the query information and the object profile information, feature annotation is performed on each of the multiple recall results to obtain annotation information for each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or the object profile information.

[0008] Based on the annotation information of each recall result, the query information, and the object profile information, the multiple recall results are sorted.

[0009] On the other hand, embodiments of this application provide a recall result sorting device, the device comprising:

[0010] The acquisition unit is used to acquire query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information;

[0011] The annotation unit is used to annotate each of the multiple recall results based on the query information and the object profile information to obtain annotation information for each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or the object profile information.

[0012] The sorting unit is used to sort the multiple recall results based on the annotation information of each recall result, the query information, and the object profile information.

[0013] In some embodiments, when the annotation unit is used to perform feature annotation on each of the plurality of recall results based on the query information and the object profile information to obtain the annotation information of each recall result, it is specifically used for:

[0014] The query information, the object profile information, and the various recall results are integrated to obtain the structured feature representations corresponding to each recall result;

[0015] Based on the structured feature representation of each recall result, feature annotation is performed on each recall result to obtain the annotation information of each recall result.

[0016] In some embodiments, when the annotation unit performs feature integration on the query information, the object profile information, and the various recall results to obtain structured feature representations of the various recall results, it is specifically used for:

[0017] The query information, the object profile information, and the various recall results are subjected to data preprocessing, which includes at least one of word segmentation, stop word removal, and syntactic analysis.

[0018] Feature extraction is performed on the query information, object profile information and each recall result after data preprocessing to obtain the corresponding query features, object profile features and each recall result feature;

[0019] The query features, the object profile features, and the features of each recall result are integrated to obtain a structured feature representation of each recall result.

[0020] In some embodiments, the annotation information of each recall result includes feature scores and natural language explanations corresponding to each preset feature among multiple preset features of each recall result. The feature scores corresponding to the preset features of the recall result are used to indicate the degree of matching between the recall result and the preset features, and the natural language explanations are used to indicate the reasons for giving the feature scores.

[0021] When the sorting unit sorts the multiple recall results based on the annotation information, the query information, and the object profile information, it is specifically used for:

[0022] The natural language interpretation vectors corresponding to each preset feature of each recall result are vectorized to obtain the natural language interpretation vectors of each preset feature.

[0023] The natural language interpretation vectors of each preset feature are concatenated to obtain the joint embedding vector of each recall result;

[0024] Based on the joint embedding vector of each recall result, the query information, and the object profile information, the multiple recall results are sorted.

[0025] In some embodiments, the preset feature includes at least one of the following:

[0026] Timeliness characteristics used to indicate the time matching degree between the recall results and the query information;

[0027] Relevance features used to indicate the text matching degree between the recall results and the query information;

[0028] Quality characteristics used to indicate the content quality of the recall results;

[0029] Personalized features used to indicate the degree of consistency between the recall results and the object profile information.

[0030] In some embodiments, the step of performing feature annotation on each of the multiple recall results based on the query information and the object profile information to obtain the annotation information of each recall result is performed by a pre-trained feature annotation model. The recall result ranking device further includes a feature annotation model training unit, which is used for:

[0031] Obtain the first object profile information sample, the first query information sample, the first recall result sample, and the label information tag corresponding to the first recall result sample of the first sample object;

[0032] Using the feature annotation model to be trained, based on the first object profile information sample and the first query information sample, annotation information is generated for the first recall result sample to obtain the predicted annotation information of the first recall result sample.

[0033] The labeling loss is determined based on the difference between the labeled information and the predicted labeled information.

[0034] The feature annotation model to be trained is obtained by training the feature annotation model based on the annotation loss.

[0035] In some embodiments, the ranking of the multiple recall results based on the annotation information, the query information, and the object profile information is performed by a pre-trained ranking model. The recall result ranking device further includes a ranking model training unit, which is used for:

[0036] Obtain the second object profile information sample, the second query information sample, and the second recall result sample of the second sample object;

[0037] Obtain the ranking score label of the second recall result sample, wherein the ranking score label is determined based on the feature weight and feature score corresponding to each preset feature among multiple preset features of the second recall result;

[0038] Obtain the labeled information sample of the second recall result sample;

[0039] The ranking model to be trained predicts the ranking score of the second recall result sample based on the second object profile information sample, the second query information sample, and the labeled information sample of the second recall result sample.

[0040] The score loss is determined based on the difference between the ranking score label and the predicted ranking score;

[0041] The ranking model is trained based on the scoring loss to obtain the ranking model.

[0042] In some embodiments, when the ranking model training unit is used to obtain the ranking score labels of the second recall result samples, it is specifically used to:

[0043] Based on the second object profile information sample and the second query information sample, multiple preset features of the second recall result sample are weighted to obtain the feature weights corresponding to each preset feature among the multiple preset features.

[0044] The feature annotation model predicts feature scores for each preset feature of the second recall result sample based on the second object profile information sample and the second query information sample, thereby obtaining the feature scores for each preset feature of the second recall result sample.

[0045] The ranking score label of the second recall result sample is determined based on the feature weights and feature scores of each preset feature of the second recall result sample.

[0046] In some embodiments, when the ranking model training unit determines the ranking score label of the second recall result sample based on the feature weights and feature scores of each preset feature of the second recall result sample, it is specifically used for:

[0047] The weight of each preset feature is multiplied by its corresponding feature score to obtain the weighted score of each preset feature.

[0048] The weighted scores of the multiple preset features are summed to obtain the ranking score label of the second recall result.

[0049] On the other hand, an embodiment of this application provides a computer-readable storage medium storing a computer program adapted for loading by a processor to execute the recall result sorting method as described in any of the above embodiments.

[0050] On the other hand, an embodiment of this application provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the recall result sorting method as described in any of the above embodiments by calling the computer program stored in the memory.

[0051] On the other hand, an embodiment of this application provides a computer program product, including computer instructions, which, when executed by a processor, implement the recall result sorting method as described in any of the above embodiments.

[0052] The embodiments of this application provide a technical solution for obtaining query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information; based on the query information and object profile information, feature annotation is performed on each of the multiple recall results to obtain annotation information for each recall result, which is used to indicate the degree of matching between each recall result and the query information and / or object profile information; based on the annotation information of each recall result, the query information, and the object profile information, the technical solution for sorting multiple recall results, by comprehensively considering the query information, annotation information, and object profile information to sort the recall results, can more accurately understand the characteristics, preferences, historical behaviors, etc. of the query object, and improve the accuracy of the recall result sorting. In addition, the annotation information clearly indicates the degree of matching between the recall result and the query information and / or object profile information, which improves the interpretability of the recall result sorting results. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a schematic diagram illustrating an application scenario for the recall result ranking system provided in this application embodiment.

[0055] Figure 2 This is a flowchart illustrating the recall result sorting method provided in an embodiment of this application.

[0056] Figure 3 This is a schematic diagram illustrating the process of weighting multiple preset features of the second recall result sample provided in an embodiment of this application.

[0057] Figure 4 This is a schematic diagram of the recall result sorting device provided in an embodiment of this application.

[0058] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0060] This application provides a method, apparatus, storage medium, device, and program product for ranking recall results. Exemplarily, the recall result ranking method of this application can be executed by a computer device, which can be a terminal or a server. The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart speaker, wearable smart device, personal computer (PC), smart vehicle terminal, etc. The terminal can also include a client, which can be a video client, shopping application client, reading application client, browser client, or instant messaging client, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0061] The embodiments of this application can be applied to scenarios such as search engines, news recommendations, artificial intelligence, e-commerce recommendations, and social media content recommendations.

[0062] First, some of the nouns or terms that appear in the description of the embodiments of this application are explained as follows:

[0063] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language, the language people use in daily life, and is closely related to linguistics; it also involves computer science and mathematics. Pre-trained models, a crucial technique for model training in artificial intelligence, evolved from large language models in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0064] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.

[0065] Large Language Model (LLM): refers to a large language model with massive parameters and training data, such as OpenAI's GPT series models, used for processing and generating natural language tasks.

[0066] Causal reasoning: A graphical representation based on causal reasoning, which performs feature weight analysis on each preset feature of the recalled document based on the object profile information of the query object, query information, and content of the recall results. The feature weight represents the strength or importance of the influence of the corresponding preset feature on the ranking of the recall results.

[0067] Jieba is an open-source toolkit focused on Chinese word segmentation, widely used for Chinese text processing tasks. It is characterized by its lightweight nature, high efficiency, and ease of use, making it particularly suitable for processing Chinese text.

[0068] NLTK is a comprehensive toolkit for English text processing. It is not only a word segmentation tool, but also includes functions such as part-of-speech tagging, syntactic analysis, and semantic analysis.

[0069] SpaCy is an industrial-grade NLP toolkit focused on efficiency and practicality. It supports multiple languages ​​(including Chinese) and offers features such as word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis.

[0070] BERT is a pre-trained language model based on the Transformer architecture. It learns rich semantic and syntactic information from text by performing unsupervised learning on large-scale text data.

[0071] Word2Vec is a technique that converts words in text into vector representations. It aims to map words to a low-dimensional real vector space, so that words with similar meanings are closer together in the vector space.

[0072] The solutions provided in this application involve technologies such as ranking recall results using artificial intelligence, which are specifically illustrated in the following embodiments. Detailed descriptions are provided below. It should be noted that the order of description in the following embodiments is not intended to limit the priority of the embodiments.

[0073] It is understood that in the specific implementation of this application, data related to the object profile data, query information, recall results, etc., of the query object are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0074] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the recall result ranking system provided in this application embodiment. The system can implement a recall result ranking method. The recall result ranking system includes a client 10, a server 20, and a network 30. The client 10 and the server 20 can interact with each other via the network 30, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.

[0075] In some embodiments, when a query object has a need for information retrieval or content recommendation, the query object inputs query information through the graphical interface of client 10. Client 10 generates a corresponding query request and sends it to server 20. For example, in a news recommendation scenario, client 10 can generate a query request based on the query currently input by the query object (such as "today's technology news") and send it to server 20. Server 20, as the core management and ranking center of the system, receives the query request from client 10, parses the request information, obtains the query information, the object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information. Server 20 can use a feature annotation model to perform feature annotation on each of the multiple recall results based on the query information and object profile information to obtain the annotation information of each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or object profile information. Server 20 can use a ranking model to rank the multiple recall results based on the annotation information of each recall result, the query information, and the object profile information, and send the ranking result to client 10.

[0076] This application provides a method for sorting recall results. This method can be executed by a terminal or a server, or by both a terminal and a server. This application uses the example of a recall result sorting method being executed by a server to illustrate the method.

[0077] Please see Figure 2 , Figure 2This is a flowchart illustrating the recall result ranking method provided in an embodiment of this application. The method may include the following steps 110 to 130:

[0078] Step 110: Obtain query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information.

[0079] The server obtains query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information.

[0080] A query is used to express the information needs of the queried object. It can be the search keywords or phrases entered by the queried object in an information retrieval system. For example, the queried object might enter "today's technology news" in a search engine or search for "smartphones" on an e-commerce platform. Query information is the basis for recall and ranking, determining which relevant content needs to be filtered from massive amounts of data.

[0081] The user profile information of the query object is a personalized description of the query object, which may include the query object's historical behavior, interests and preferences, age, gender, geographical location, etc., such as the query object's historical search records, click behavior, purchase records, etc.

[0082] The multiple recall results (docs) are preliminary results selected based on the query information. They may contain a lot of information, but have not yet undergone fine-grained sorting.

[0083] This application's embodiments introduce object profile information and annotation information, and sort the recall results by analyzing the historical behavioral data of the query object (such as clicks, browsing history, etc.). This personalized sorting method ensures that the final ranking better matches the interests and needs of the specific query object, improving the search experience and satisfaction.

[0084] Step 120: Based on the query information and the object profile information, feature annotation is performed on each of the multiple recall results to obtain the annotation information of each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or the object profile information.

[0085] Based on query information and object profile information, the server performs feature annotation on each recall result in multiple recall results to obtain annotation information for each recall result.

[0086] In some embodiments, the server can use a feature annotation model to annotate each of the multiple recall results based on query information and object profile information, thereby obtaining the annotation information of each recall result.

[0087] The feature annotation model may be a large language model, such as GPT-4 of OpenAI or higher versions.

[0088] In some embodiments, in step 120, performing feature annotation on each of a plurality of recall results based on query information and object portrait information to obtain annotation information of each recall result may include the following steps 1201 to 1202:

[0089] Step 1201: performing feature integration on the query information, the object portrait information and each recall result to obtain a structured feature representation corresponding to each recall result.

[0090] For each recall result doc among the plurality of recall results, a server performs feature integration on the query information query, the object portrait information user and each recall result doc to obtain a structured feature representation user-query-doc corresponding to each recall result, and the structured feature representation includes the feature of the query information query, the object portrait information user and the document feature of the corresponding recall result doc.

[0091] The data type of the structured feature representation user-query-doc of the recall result may be a semantic vector.

[0092] In some embodiments, in step 1201, performing feature integration on the query information, the object portrait information and each recall result to obtain the structured feature representation corresponding to each recall result may include the following steps 12011 to 12013:

[0093] Step 12011: performing data preprocessing on the query information, the object portrait information and each recall result, wherein the data preprocessing includes at least one of word segmentation, stop word removal and grammatical analysis.

[0094] Wherein, for a piece of text information, word segmentation refers to splitting the text information into individual independent words; stop word removal refers to removing common words that make no substantial contribution to semantic expression, such as "of", "is", "in" and the like; grammatical analysis is used to understand the grammatical structure of the text information, so as to extract features better.

[0095] For example, for the query information "today's technology news", a word segmentation tool (such as Jieba, NLTK, SpaCy, etc.) can be used to segment the query information "today's technology news" into ["today", "'s", "technology", "news"].

[0096] When removing stop words, a stop word list (such as a Chinese stop word list or an English stop word list) is loaded, and stop words in the word segmentation result are removed. For example, for the query information "today's technology news", after removing the stop word "'s", the word segmentation result becomes ["today", "technology", "news"].

[0097] A syntactic analysis tool (such as Stanford NLP, SpaCy, etc.) is used to perform part-of-speech tagging, dependency syntactic analysis and other processing on the query information, object portrait information and each recall result.

[0098] For example, for the query information "today's technology news", it is analyzed that "technology" is a noun and "today" is a temporal adverbial.

[0099] In step 12012, feature extraction is performed on the preprocessed query information, object portrait information and each recall result to obtain corresponding query features, object portrait features and each recall result feature.

[0100] Based on the data preprocessing, the server performs feature extraction on the preprocessed query information, object portrait information and each recall result, so as to dig out representative features from the original information including the preprocessed query information, object portrait information and each recall result, and obtain corresponding query features, object portrait features and each recall result feature for subsequent model processing.

[0101] For example, keywords in the query information, interest tags in the object portrait information, and titles, abstracts, time stamps, type tags and the like in the recall results are extracted.

[0102] In step 12013, feature integration processing is performed on the query features, object portrait features and each recall result feature to obtain structured feature representations of each recall result.

[0103] The server performs feature integration processing on the query features, object portrait features and each recall result feature obtained after feature extraction, to obtain structured feature representations of each recall result.

[0104] In some embodiments, the server inputs the query features, object portrait features and each recall result feature obtained after feature extraction into a pre-trained semantic vector model for feature integration processing, so as to obtain structured feature representations of each recall result. The pre-trained semantic vector model is configured to convert text information into vector representations with semantic information, so that the features of the query information, object portrait information and recall results are more structured and easier to process.

[0105] Wherein, the pre-trained semantic vector model can be BERT, Word2Vec, GPT or other models, and the present application does not impose restrictions on the specific model structure of the pre-trained semantic vector model.

[0106] Step 1202: Based on the structured feature representation corresponding to each recall result, perform feature annotation on each recall result to obtain the annotation information of each recall result.

[0107] For each recall result among multiple recall results, the server performs feature annotation on the recall result based on the structured features of each recall result, and obtains the annotation information of each recall result.

[0108] In some embodiments, the server can generate annotation information for the recall results using a feature annotation model. Specifically, the server inputs the structured feature representation corresponding to each recall result into the feature annotation model to obtain the annotation information for each recall result.

[0109] For example, for the query "today's technology news", the corresponding output of the feature labeling model is: "The feature score of the relevance feature is 0.8, and the natural language interpretation of the relevance feature is: This recall result is highly relevant to technology news; the feature score of the timeliness feature is 0.9, and the natural language interpretation of the timeliness feature is: This recall result was published today; the feature score of the quality feature is 0.9, and the natural language interpretation of the quality feature is: This recall result comes from a well-known technology website; the feature score of the personalization feature is 0.8, and the natural language interpretation of the personalization feature is: This recall result matches the historical preferences of the query object."

[0110] Step 130: Sort the multiple recall results based on the annotation information, query information and object profile information of each recall result.

[0111] The server sorts multiple recall results based on the annotation information, query information, and object profile information of each recall result.

[0112] In some embodiments, the ranking score corresponding to each recall result can be calculated based on the annotation information, query information and object profile information of each recall result, and then the multiple recall results can be ranked according to the ranking score corresponding to each recall result.

[0113] In some embodiments, the server can rank multiple recall results based on the annotation information, query information, and object profile information of each recall result using a ranking model.

[0114] In some embodiments, the ranking model first calculates the ranking score corresponding to each recall result, and then ranks the multiple recall results according to the ranking score corresponding to each recall result.

[0115] In some embodiments, the annotation information of each recall result includes the feature score and natural language explanation corresponding to each preset feature among multiple preset features of each recall result. The feature score corresponding to the preset feature of the recall result is used to indicate the degree of matching between the recall result and the preset feature, and the natural language explanation is used to indicate the reason for giving the feature score.

[0116] In some embodiments, the preset features include at least one of the following:

[0117] A time-sensitive feature used to indicate the time matching degree between recall results and query information;

[0118] A relevance feature used to indicate the textual matching degree between the recall results and the query information;

[0119] Quality characteristics used to indicate the content quality of recall results;

[0120] Personalized features used to indicate the degree of consistency between recall results and object profile information.

[0121] Specifically, the timeliness feature refers to the degree of match between the publication or update time of the recall result and the timeliness requirement of the query information. The higher the feature score of the timeliness feature of the recall result, the higher the time matching degree between the recall result and the query information. For example, if the query information is "today's technology news", the closer the publication date of the recall result is to the current time, the higher the feature score of the timeliness feature.

[0122] Relevance features are used to indicate the textual match between the retrieved results and the query information. Specifically, relevance features measure the semantic relevance between the content of the retrieved results and the query information. The higher the feature score of the relevance feature, the higher the textual match between the retrieved results and the query information. For example, if the query information is "the latest advances in artificial intelligence," then if the retrieved results contain detailed discussions or the latest research results related to artificial intelligence, their relevance feature score will be higher.

[0123] Quality features are used to indicate the content quality of the recall results, including aspects such as accuracy, completeness, authority, readability, and credibility. A higher feature score indicates higher content quality. For example, if the recall results come from authoritative sources (such as well-known journals or articles written by experts) and are detailed and logically clear, their quality feature scores will be higher.

[0124] Personalized features are used to indicate the degree of consistency between the recall results and the target user profile information. Target user profile information can include a user's interests, preferences, historical behavior, etc. A higher feature score for a personalized feature indicates a higher degree of consistency between the recall results and the target user profile information. For example, if a user has a historical preference for content related to "machine learning," then if the recall results include machine learning-related content, the feature score for their personalized feature will be higher.

[0125] The feature score and natural language interpretation of each preset feature work together to help understand why a certain recall result is recommended and to evaluate whether the recall result meets the requirements.

[0126] In some embodiments, in step 130, sorting multiple recall results based on the annotation information, query information, and object profile information of each recall result may include the following steps 1301 to 1303:

[0127] Step 1301: Vectorize the natural language interpretation of each preset feature of each recall result to obtain the natural language interpretation vector of each preset feature.

[0128] For each preset feature of each recall result, the server uses a natural language processing model to convert the natural language interpretation of the preset feature into a fixed-dimensional vector representation, thus obtaining the natural language vector of the preset feature.

[0129] Specifically, for each preset feature of each recall result, the natural language explanation of the preset feature is input into a pre-trained natural language processing model, and the output vector of the natural language processing model is extracted to obtain the natural language explanation vector for each preset feature. Thus, for each recall result, the natural language explanation corresponding to each preset feature of the recall result is converted into a natural language explanation vector.

[0130] In some embodiments, the pre-trained natural language processing model can be the Sentence-BERT model. Sentence-BERT (SBERT) is a model developed based on BERT specifically for processing sentence-level semantic understanding.

[0131] Step 1302: Concatenate the natural language interpretation vectors of each preset feature to obtain the joint embedding vector of each recall result.

[0132] For each recall result, the natural language interpretation vectors of all its preset features are concatenated in order to form a joint embedding vector. Each recall result corresponds to a joint embedding vector, which integrates all the natural language interpretations corresponding to all the preset features of the recall result.

[0133] In some embodiments, if the dimensions of the natural language interpretation vectors corresponding to different preset features are inconsistent, each natural language interpretation vector can be mapped to a unified dimension through a fully connected layer before being concatenated.

[0134] In some embodiments, steps 1301 to 1302 may be performed by the annotation reason embedding module.

[0135] This embodiment transforms the natural language explanation (or scoring reason) corresponding to each preset feature into a joint embedding vector. This joint embedding vector is added as an additional feature to the ranking of the recall results. In this way, the ranking model can not only output the ranking score of the recall results but also provide the scoring reason, making the scoring process more interpretable. Furthermore, the joint embedding vector helps expand the feature space, enabling the ranking model to better capture subtle differences between the content of the recall results and the needs of the query object, thereby further improving scoring accuracy.

[0136] Step 1303: Sort the multiple recall results based on the joint embedding vector, query information and object profile information of each recall result.

[0137] In some embodiments, the server sorts multiple recall results based on the joint embedding vector of each recall result, query information, and object profile information using a ranking model.

[0138] The solution provided in this embodiment, by combining natural language interpretation of the preset features of the recall results, query information, and object profile information, can more comprehensively capture the timeliness, relevance, quality, and personalization needs of the recall results, thereby improving the accuracy of ranking and user satisfaction.

[0139] In some embodiments, the multiple recall results are sorted based on the joint embedding vector of the annotation information of each recall result, query information, and object profile information, including:

[0140] Data preprocessing is performed on query information, target profile information, and each recall result. Data preprocessing includes at least one of the following: word segmentation, stop word removal, and syntactic analysis.

[0141] Feature extraction is performed on the query information, object profile information, and various recall results after data preprocessing.

[0142] The query information after feature extraction, the object profile information, and each recall result are input into the pre-trained semantic vector model to obtain the structured feature representation of each recall result;

[0143] The joint embedding vector of the structured feature representation and annotation information of each recall result is input into the ranking model to rank the multiple recall results.

[0144] In some embodiments, feature annotation is performed on each of the multiple recall results based on query information and object profile information to obtain the annotation information of each recall result. This is performed by a pre-trained feature annotation model. The training steps of the feature annotation model may include the following steps 11 to 14:

[0145] Step 11: Obtain the first object profile information sample, the first query information sample, the first recall result sample, and the label information tag corresponding to the first recall result sample of the first sample object.

[0146] The server collects the first sample data, including the first object profile information sample of the first sample object, the first query information sample, the first recall result sample, and the annotation information tags corresponding to the first recall result sample.

[0147] In some embodiments, each first sample data can be a quadruple: (first object profile information sample, first query information sample, first recall result sample, annotation information label corresponding to the first recall result sample).

[0148] In some embodiments, the label information corresponding to the first recall result sample can be obtained from the search behavior of the first sample object. Specifically, it can be generated using rules designed based on business logic and domain knowledge, or it can be obtained through manual labeling.

[0149] Step 12: Using the feature annotation model to be trained, based on the first object profile information sample and the first query information sample, generate annotation information for the first recall result sample to obtain the predicted annotation information of the first recall result sample.

[0150] The first object profile information sample, the first query information sample, and the first recall result sample, which have undergone data preprocessing (e.g., word segmentation, stop word removal, and syntactic analysis) and feature extraction, are input into the feature annotation model to be trained. The feature annotation model to be trained outputs the predicted annotation information of the first recall result.

[0151] Step 13: Determine the annotation loss based on the difference between the labeled information and the predicted annotation information.

[0152] The server uses a pre-defined labeling loss function to determine the labeling loss based on the difference between the labeled information and the predicted labeling information.

[0153] In some embodiments, the annotation information labels include feature score labels and natural language interpretation labels corresponding to each of the multiple preset features, and the annotation loss includes feature score loss and natural language interpretation loss. Then, when training the feature annotation model, a multi-task training objective can be constructed, including a feature score prediction task and a natural language interpretation generation task.

[0154] Specifically, for the predicted feature scores in the predicted annotation information, the mean squared error or cross-entropy loss function can be used to calculate the feature score loss between the feature score label and the predicted feature score. For the predicted natural language interpretation in the predicted annotation information, the loss function of the text generation task (such as cross-entropy loss) is used to calculate the natural language interpretation loss between the natural language interpretation label and the predicted natural language interpretation. The feature score loss and the natural language interpretation loss are weighted and summed to obtain the annotation loss.

[0155] Step 14: Train the feature annotation model to be trained based on the annotation loss to obtain the feature annotation model.

[0156] The server trains the feature annotation model to be trained based on the annotation loss, and obtains the feature annotation model.

[0157] In some embodiments, training the feature labeling model to be trained based on the labeling loss includes: calculating the gradient of the labeling loss function on the model parameters of the feature labeling model to be trained, and updating the model parameters of the feature labeling model to be trained using gradient descent, for example, using optimizers such as Adam or SGD for parameter updating.

[0158] Repeat steps 11 to 14 until the feature annotation model converges or reaches the predetermined number of training rounds.

[0159] In some embodiments, ranking multiple recall results based on the annotation information, query information, and object profile information of each recall result is performed by a pre-trained ranking model. The training steps of the ranking model may include the following steps 21 to 26:

[0160] Step 21: Obtain the second object profile information sample, the second query information sample, and the second recall result sample of the second sample object.

[0161] Step 22: Obtain the ranking score label of the second recall result sample. The ranking score label is determined based on the feature weight and feature score of each preset feature among the multiple preset features of the second recall result.

[0162] The server retrieves the sorting score labels of the second recall result samples.

[0163] In some embodiments, obtaining the ranking score label of the second recall result sample may include the following steps 221 to 223:

[0164] Step 221: Based on the second object profile information sample and the second query information sample, assign weights to multiple preset features of the second recall result sample to obtain the feature weights corresponding to each preset feature among the multiple preset features.

[0165] In some embodiments, a causal relationship graph generation model can be used to determine the causal relationship graph corresponding to each second recall result sample based on the second object profile information sample and the second query information sample. The causal relationship graph includes the feature weights of each preset feature.

[0166] In some embodiments, the causal relationship graph generation model can be a large language model.

[0167] In some embodiments, the causal relationship graph generation model can be a graph neural network model.

[0168] Specifically, when using a graph neural network (Graph Neural Network) model as a causal relationship graph generation model, information such as the second query information sample, the second recall result sample, and the second object profile information sample can be regarded as graph nodes. The relationships and weights between nodes are dynamically learned and updated through the propagation mechanism of the Graph Neural Network. The advantage of Graph Neural Networks lies in their ability to adaptively extract structured information from large-scale data and automatically optimize the weight allocation of each dimension through the graph propagation mechanism. This method effectively avoids manually setting causal relationships while improving the model's performance on large-scale data.

[0169] Causal graph generation models based on graph neural networks possess strong expressive power, enabling them to more accurately capture complex relationships between features, and are particularly suitable for processing large-scale, structured datasets. Furthermore, the advantages of graph neural networks in their graph structure give them greater flexibility and robustness when handling complex relationships.

[0170] In some embodiments, a causal relationship graph generation model is used to determine the causal relationship graph corresponding to each second recall result sample based on the second object profile information sample and the second query information sample, including:

[0171] The causal relationship graph model first determines the query demand type of the second query information sample. The query demand types include time-sensitive queries, professional search queries, personal interest queries, and general search queries.

[0172] The causal relationship graph model determines the weight relationship of each preset feature of the second recall result sample based on the query demand type of the second query information sample.

[0173] The causal relationship graph model determines the feature weights of each preset feature in the second recall result sample based on the weight relationship of each preset feature.

[0174] Please see Figure 3 , Figure 3 A flowchart illustrating the weight allocation of multiple preset features for the second recall result samples provided in this application embodiment:

[0175] First, determine whether the query requirement type of the second query information sample is a time-sensitive query. If it is a time-sensitive query, then determine the weight relationship of each preset feature of the second recall result sample as follows: the feature weight of time-sensitive feature > the feature weight of relevance feature > the feature weight of quality feature > the feature weight of personalized feature.

[0176] If it is not a time-sensitive query, determine whether it is a professional search query. If it is a professional search query, determine the weight relationship of each preset feature of the second recall result sample as follows: feature weight of quality feature > feature weight of relevance feature > feature weight of time-sensitive feature > feature weight of personalized feature.

[0177] If it is not a professional search query, determine whether it is a personal interest query. If it is a personal interest query, determine the weight relationship of each preset feature of the second recall result sample as follows: feature weight of personalized feature > feature weight of quality feature > feature weight of relevance feature > feature weight of timeliness feature.

[0178] If the query is not for personal interest, then the query requirement type of the second query information sample can be determined to be a general search query. Therefore, the weight relationship of each preset feature of the second recall result sample is determined as follows: feature weight of relevance feature > feature weight of quality feature > feature weight of timeliness feature > feature weight of personalization feature.

[0179] Analyze the characteristics of the second object profile information sample and the second recall result sample;

[0180] The feature weight of the timeliness feature is determined based on the time matching degree between the release time or update time of the second recall result sample and the time of the second query information sample.

[0181] The feature weights of the relevance features are determined based on the text matching degree between the content of the second recall result sample and the second query information sample.

[0182] The feature weights of quality characteristics are determined based on the credibility, authority, and accuracy of the second recall results samples.

[0183] The feature weights of personalized features are determined based on the degree of consistency between the second recall result sample and the second object profile information sample.

[0184] After determining the relative weights of each preset feature, the feature weights can be fine-tuned based on the target audience profile and recall results, building upon the default weights. For example, for time-sensitive queries, the default weight for timeliness is higher, but if the target audience has strong personalized preferences, the weight for personalization can be increased appropriately.

[0185] For example, for the second query information sample "Today's technology news", it can be determined that it belongs to the timeliness query, because the second query object sample obviously wants to see the latest news content. Therefore, it can be determined that the weight relationship of each preset feature of the second recall result sample is: the feature weight of timeliness feature > the feature weight of relevance feature > the feature weight of quality feature > the feature weight of personalization feature.

[0186] For the three secondary recall results samples (News 1, News 2, and News 3), further analysis is needed to determine the feature weights of each preset feature for each secondary recall result sample:

[0187] News 1 was published 2 hours ago, so the feature weight for the timeliness feature of News 1 can be 0.4. The content of News 1 is highly matched with the query "technology news", so the feature weight for the relevance feature of News 1 can be 0.3. News 1 comes from a well-known technology website, so the feature weight for the quality feature can be 0.2. News 1 matches the user's historical preferences, so the feature weight for the personalization feature can be 0.1.

[0188] News 2 was published 1 day ago, so the feature weight for the timeliness feature of News 1 can be 0.4. The content of News 2 is relatively relevant to the query "technology news", but it involves other fields, so the feature weight for the relevance feature of News 2 can be 0.3. News 2 comes from a secondary technology website, so the feature weight for the quality feature can be 0.2. News 2 matches some of the user's historical preferences, so the feature weight for the personalization feature can be 0.1.

[0189] News 3 was published 3 days ago, so the feature weight for the timeliness feature of News 1 can be 0.4. The content of News 3 is relatively broad and does not fully conform to the theme of "technology", so the feature weight for the relevance feature of News 3 can be 0.3. News 3 comes from a general news website, so the feature weight for the quality feature can be 0.2. News 3 does not conform to the user's historical preferences, so the feature weight for the personalization feature can be 0.1.

[0190] Step 222: Based on the second object profile information sample and the second query information sample, the feature annotation model is used to predict the feature scores of each preset feature of the second recall result sample, so as to obtain the feature scores of each preset feature of the second recall result sample.

[0191] In some embodiments, the server inputs a second object profile information sample, a second query information sample, and a second recall result sample that have undergone data preprocessing (e.g., word segmentation, stop word removal, and syntactic analysis) and feature extraction into the feature annotation model, and the feature annotation model outputs feature scores for each preset feature of the second recall result.

[0192] For example, if news 1 was published 2 hours ago, the feature score for the timeliness feature of news 1 can be 0.9; if the content of news 1 is highly matched with the query "technology news", the feature score for the relevance feature of news 1 can be 0.95; if news 1 comes from a well-known technology website, the feature score for the quality feature can be 0.9; and if news 1 meets the historical preference of the second query object sample for technology news, the feature score for the personalization feature can be 0.8.

[0193] News 2 was published 1 day ago, so the timeliness feature score of News 1 can be 0.7. The content of News 2 is relatively relevant to the query "technology news", but it involves other fields, so the relevance feature score of News 2 can be 0.85. News 2 comes from a secondary technology website, so the quality feature score can be 0.7. News 2 partially matches the historical preference of the second query object sample for technology news, so the personalization feature score can be 0.6.

[0194] News 3 was published 3 days ago. The timeliness feature score of News 1 can be 0.4. The content of News 3 is relatively broad and does not fully conform to the theme of "science and technology". The relevance feature score of News 3 can be 0.6. News 3 comes from a general news website. The quality feature score can be 0.6. News 3 does not conform to the historical preference of the second query object sample for science and technology news. The personalization feature score can be 0.4.

[0195] In ranking systems of related technologies, the feature weights of each feature are usually fixed, lacking flexibility and unable to be adjusted according to the specific needs of the query information. However, the embodiments of this application, through causal reasoning graphs and a large-model-based reasoning mechanism, can dynamically adjust the feature weights of each feature based on changes in the query information and the content of the recall results. This flexibility allows the ranking system to adaptively respond to different scenarios, thereby improving the accuracy and relevance of the recall result ranking.

[0196] Step 223: Determine the ranking score label of the second recall result sample based on the feature weights and feature scores of each preset feature of the second recall result sample.

[0197] The server determines the ranking score label of the second recall result sample based on the feature weights and feature scores of each preset feature of the second recall result sample.

[0198] In some embodiments, determining the ranking score label of the second recall result sample based on the feature weights and feature scores of each preset feature of the second recall result sample may include the following steps 2231 to 2232:

[0199] Step 2231: Multiply the feature weight of each preset feature by its corresponding feature score to obtain the weighted score of each preset feature.

[0200] Step 2232: Add the weighted scores of multiple preset features to obtain the ranking score label of the second recall result.

[0201] Specifically, the ranking rating labels can be calculated using a formula: Among them, w i The feature weights of the i-th preset feature are s i The feature score is given to the i-th preset feature.

[0202] For example, News 1, News 2, and News 3 in the examples above:

[0203] The ranking score for News 1 is: 0.9×0.4+0.95×0.3+0.9×0.2+0.8×0.1=0.905.

[0204] The ranking score for News 2 is: 0.7×0.4+0.85×0.3+0.7×0.2+0.6×0.1=0.735.

[0205] The ranking score for News 3 is: 0.4×0.4+0.6×0.3+0.6×0.2+0.4×0.1=0.48.

[0206] Step 23: Obtain the labeled information sample of the second recall result sample.

[0207] In some embodiments, the annotation information sample of the second recall result sample that is manually annotated can be obtained directly.

[0208] In some embodiments, the annotation information sample of the second recall result sample can be determined by the feature annotation model based on the second object profile information sample and the second query information sample.

[0209] Step 24: Based on the labeled information samples of the second object profile information sample, the second query information sample, and the second recall result sample, the ranking model to be trained performs ranking score prediction on the second recall result sample to obtain the predicted ranking score of the second recall result sample.

[0210] In some embodiments, the labeled information samples of the second object profile information sample, the second query information sample, and the second recall result sample after data preprocessing and feature extraction are input into the ranking model to be trained, and the predicted ranking score of the second recall result sample output by the ranking model to be trained is obtained.

[0211] Step 25: Determine the rating loss based on the difference between the ranked rating label and the predicted ranked rating.

[0212] In some embodiments, mean squared error or a ranking task can be used as the rating loss function to determine the rating loss based on the difference between the ranked rating label and the predicted ranked rating.

[0213] Step 26: Train the ranking model to be trained based on the rating loss to obtain the ranking model.

[0214] In some embodiments, training the ranking model to be trained based on the rating loss includes: calculating the gradient of the rating loss function on the model parameters of the ranking model to be trained, and updating the model parameters of the ranking model to be trained using gradient descent, for example, using optimizers such as Adam or SGD.

[0215] Repeat steps 21 to 26 until the ranking model converges or reaches the predetermined number of training rounds.

[0216] In related technologies, methods for training ranking models based on manually labeled ranking scores are heavily influenced by subjective factors, easily leading to inconsistencies and instability in the ranking scores. This application's embodiments, when determining ranking scores, utilize a large model to automatically generate a causal relationship graph. This allows for objective analysis of the multi-dimensional relationships between recall results and query information, avoiding the subjectivity of manual labeling, and automatically deriving the causal relationships between each preset feature (e.g., relevance, quality, timeliness, personalization). The causal graph has high interpretability, clearly showing how each preset feature interacts and influences, providing more reliable data support for ranking scores.

[0217] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.

[0218] The embodiments of this application provide a technical solution for obtaining query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information; based on the query information and object profile information, feature annotation is performed on each of the multiple recall results to obtain annotation information for each recall result, which is used to indicate the degree of matching between each recall result and the query information and / or object profile information; based on the annotation information of each recall result, the query information, and the object profile information, the technical solution for sorting multiple recall results, by comprehensively considering the query information, annotation information, and object profile information to sort the recall results, can more accurately understand the characteristics, preferences, historical behaviors, etc. of the query object, and improve the accuracy of the recall result sorting. In addition, the annotation information clearly indicates the degree of matching between the recall result and the query information and / or object profile information, which improves the interpretability of the recall result sorting results.

[0219] To facilitate better implementation of the recall result sorting method of this application embodiments, this application embodiment also provides a recall result sorting apparatus. Please refer to... Figure 4 , Figure 4 This is a schematic diagram of the recall result sorting device provided in an embodiment of this application. The recall result sorting device 400 may include:

[0220] The acquisition unit 410 is used to acquire query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information;

[0221] The annotation unit 420 is used to annotate the features of each recall result in multiple recall results based on query information and object profile information, and to obtain the annotation information of each recall result. The annotation information is used to indicate the degree of matching between each recall result and query information and / or object profile information.

[0222] The sorting unit 430 is used to sort multiple recall results based on the annotation information, query information and object profile information of each recall result.

[0223] In some embodiments, when the annotation unit 420 is used to perform feature annotation on each recall result among multiple recall results based on query information and object profile information using a feature annotation model to obtain annotation information for each recall result, it is specifically used for:

[0224] The query information, object profile information, and each recall result are integrated to obtain the structured feature representation corresponding to each recall result;

[0225] Based on the structured feature representation of each recall result, feature annotation is performed on each recall result to obtain the annotation information of each recall result.

[0226] In some embodiments, when the annotation unit 420 performs feature integration on the query information, object profile information, and each recall result to obtain a structured feature representation of each recall result, it is specifically used for:

[0227] Data preprocessing is performed on query information, object profile information and each recall result. Data preprocessing includes at least one of the following: word segmentation, stop word removal and syntactic analysis.

[0228] Feature extraction is performed on the query information, object profile information and each recall result after data preprocessing to obtain the corresponding query features, object profile features and each recall result feature;

[0229] The query features, object profile features, and features of each recall result are integrated to obtain a structured feature representation of each recall result.

[0230] In some embodiments, the annotation information of each recall result includes the feature score and natural language explanation corresponding to each preset feature among multiple preset features of each recall result. The feature score corresponding to the preset feature of the recall result is used to indicate the degree of matching between the recall result and the preset feature, and the natural language explanation is used to indicate the reason for giving the feature score.

[0231] When sorting multiple recall results based on the annotation information, query information, and object profile information of each recall result, the sorting unit 430 is specifically used for:

[0232] Vectorize the natural language interpretations corresponding to each preset feature of each recall result to obtain the natural language interpretation vectors of each preset feature.

[0233] The natural language interpretation vectors of each preset feature are concatenated to obtain the joint embedding vector of each recall result;

[0234] Based on the joint embedding vectors of each recall result, query information, and object profile information, the multiple recall results are ranked.

[0235] In some embodiments, the preset features include at least one of the following:

[0236] A time-sensitive feature used to indicate the time matching degree between recall results and query information;

[0237] A relevance feature used to indicate the textual matching degree between the recall results and the query information;

[0238] Quality characteristics used to indicate the content quality of recall results;

[0239] Personalized features used to indicate the degree of consistency between recall results and object profile information.

[0240] In some embodiments, feature annotation is performed on each recall result among multiple recall results based on query information and object profile information. The annotation information of each recall result is obtained by a pre-trained feature annotation model. The recall result ranking device further includes a feature annotation model training unit, which is used for:

[0241] Obtain the first object profile information sample, the first query information sample, the first recall result sample, and the label information tag corresponding to the first recall result sample of the first sample object;

[0242] Using the feature annotation model to be trained, based on the first object profile information sample and the first query information sample, the annotation information of the first recall result sample is generated to obtain the predicted annotation information of the first recall result sample.

[0243] The annotation loss is determined based on the difference between the labeled information and the predicted labeled information;

[0244] The feature annotation model is trained using the annotation loss method to obtain the feature annotation model.

[0245] In some embodiments, ranking multiple recall results based on the annotation information, query information, and object profile information of each recall result is performed by a pre-trained ranking model. The recall result ranking device further includes a ranking model training unit, which is used for:

[0246] Obtain the second object profile information sample, the second query information sample, and the second recall result sample of the second sample object;

[0247] Obtain the ranking score label of the second recall result sample. The ranking score label is determined based on the feature weight and feature score of each preset feature among multiple preset features of the second recall result.

[0248] Obtain the labeled information sample of the second recall result sample;

[0249] The ranking model to be trained predicts the ranking score of the second recall result sample based on the labeled information samples of the second object profile information sample, the second query information sample, and the second recall result sample.

[0250] The score loss is determined based on the difference between the ranked score labels and the predicted ranked scores.

[0251] The ranking model is obtained by training the ranking model to be trained based on the scoring loss.

[0252] In some embodiments, when the ranking model training unit is used to obtain the ranking score labels for the second recall result samples, it is specifically used to:

[0253] Based on the second object profile information sample and the second query information sample, multiple preset features of the second recall result sample are weighted to obtain the feature weights corresponding to each preset feature among the multiple preset features.

[0254] The feature annotation model predicts the feature scores of each preset feature of the second recall result sample based on the second object profile information sample and the second query information sample, and obtains the feature scores of each preset feature of the second recall result sample.

[0255] The ranking score label of the second recall result sample is determined based on the feature weights and feature scores of each preset feature of the second recall result sample.

[0256] In some embodiments, when the ranking model training unit determines the ranking score label of the second recall result sample based on the feature weights and feature scores of each preset feature of the second recall result sample, it is specifically used for:

[0257] The weight of each preset feature is multiplied by its corresponding feature score to obtain the weighted score of each preset feature.

[0258] The weighted scores of multiple preset features are summed to obtain the ranking score label of the second recall result.

[0259] It should be noted that the functions of each module in the recall result sorting device 400 in this application embodiment can be referred to the specific implementation of any embodiment in the above method embodiments, and will not be repeated here.

[0260] Each unit in the above-described device can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.

[0261] For example, the recall result sorting device 400 can be integrated into a terminal or server that has storage and a processor and thus computing power, or the recall result sorting device 400 can be the terminal or server.

[0262] In some embodiments, this application also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0263] Figure 5 A schematic structural diagram of the computer device provided in the embodiments of this application, such as Figure 5As shown, the computer device 500 may include: a communication interface 501, a memory 502, a processor 503, and a communication bus 504. The communication interface 501, memory 502, and processor 503 communicate with each other via the communication bus 504. The communication interface 501 is used for data communication between the device 500 and external devices. The memory 502 can be used to store software programs and modules, and the processor 503 runs the software programs and modules stored in the memory 502, such as the software programs for the corresponding operations in the aforementioned method embodiments.

[0264] In some embodiments, the processor 503 may invoke software programs and modules stored in the memory 502 to perform the following operations: obtain query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information; based on the query information and object profile information, perform feature annotation on each of the multiple recall results to obtain annotation information for each recall result, the annotation information being used to indicate the degree of matching between each recall result and the query information and / or object profile information; and sort the multiple recall results based on the annotation information of each recall result, the query information, and the object profile information.

[0265] In some embodiments, the computer device 500 may be integrated into a terminal or server that has storage and a processor, thus possessing computing capabilities; or the computer device 500 may be such a terminal or server. The terminal may be a smartphone, tablet, laptop, smart TV, smart speaker, wearable smart device, personal computer, or other similar device. The server may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0266] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the methods described above in the embodiments of this application; for brevity, further details are omitted here.

[0267] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0268] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0269] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0270] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0271] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0272] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0273] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0274] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0275] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0276] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0277] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0278] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for ranking recall results, characterized in that, The method includes: Obtain query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information; Based on the query information and the object profile information, feature annotation is performed on each of the multiple recall results to obtain annotation information for each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or the object profile information. Based on the annotation information of each recall result, the query information, and the object profile information, the multiple recall results are sorted.

2. The recall result ranking method as described in claim 1, characterized in that, Based on the query information and the object profile information, feature annotation is performed on each of the multiple recall results to obtain the annotation information of each recall result, including: The query information, the object profile information, and the various recall results are integrated to obtain the structured feature representations corresponding to each recall result; Based on the structured feature representation of each recall result, feature annotation is performed on each recall result to obtain the annotation information of each recall result.

3. The recall result ranking method as described in claim 2, characterized in that, The step of integrating the query information, the object profile information, and the various recall results to obtain the structured feature representation of each recall result includes: The query information, the object profile information, and the various recall results are subjected to data preprocessing, which includes at least one of word segmentation, stop word removal, and syntactic analysis. Feature extraction is performed on the query information, object profile information and each recall result after data preprocessing to obtain the corresponding query features, object profile features and each recall result feature; The query features, the object profile features, and the features of each recall result are integrated to obtain a structured feature representation of each recall result.

4. The recall result ranking method as described in claim 1, characterized in that, The annotation information of each recall result includes the feature score and natural language explanation corresponding to each preset feature among the multiple preset features of each recall result. The feature score corresponding to the preset feature of the recall result is used to indicate the degree of matching between the recall result and the preset feature, and the natural language explanation is used to indicate the reason for giving the feature score. Based on the annotation information of each recall result, the query information, and the object profile information, the multiple recall results are sorted, including: The natural language interpretation vectors corresponding to each preset feature of each recall result are vectorized to obtain the natural language interpretation vectors of each preset feature. The natural language interpretation vectors corresponding to each preset feature are concatenated to obtain the joint embedding vector of each recall result; Based on the joint embedding vector of each recall result, the query information, and the object profile information, the multiple recall results are sorted.

5. The recall result ranking method as described in claim 4, characterized in that, The preset feature includes at least one of the following: Timeliness characteristics used to indicate the time matching degree between the recall results and the query information; Relevance features used to indicate the text matching degree between the recall results and the query information; Quality characteristics used to indicate the content quality of the recall results; Personalized features used to indicate the degree of consistency between the recall results and the object profile information.

6. The recall result ranking method as described in claim 1, characterized in that, The step of performing feature annotation on each of the multiple recall results based on the query information and the object profile information to obtain the annotation information of each recall result is performed by a pre-trained feature annotation model. The training steps of the feature annotation model include: Obtain the first object profile information sample, the first query information sample, the first recall result sample, and the label information tag corresponding to the first recall result sample of the first sample object; Using the feature annotation model to be trained, based on the first object profile information sample and the first query information sample, annotation information is generated for the first recall result sample to obtain the predicted annotation information of the first recall result sample. The labeling loss is determined based on the difference between the labeled information and the predicted labeled information. The feature annotation model to be trained is obtained by training the feature annotation model based on the annotation loss.

7. The recall result ranking method as described in claim 1, characterized in that, The ranking of the multiple recall results based on the annotation information, query information, and object profile information is performed by a pre-trained ranking model. The training steps of the ranking model include: Obtain the second object profile information sample, the second query information sample, and the second recall result sample of the second sample object; Obtain the ranking score label of the second recall result sample, wherein the ranking score label is determined based on the feature weight and feature score corresponding to each preset feature among multiple preset features of the second recall result; Obtain the labeled information sample of the second recall result sample; The ranking model to be trained predicts the ranking score of the second recall result sample based on the second object profile information sample, the second query information sample, and the labeled information sample of the second recall result sample. The score loss is determined based on the difference between the ranking score label and the predicted ranking score; The ranking model is trained based on the scoring loss to obtain the ranking model.

8. The recall result ranking method as described in claim 7, characterized in that, Obtain the ranking and scoring labels of the second recall result samples, including: Based on the second object profile information sample and the second query information sample, multiple preset features of the second recall result sample are weighted to obtain the feature weights corresponding to each preset feature among the multiple preset features. The feature annotation model predicts feature scores for each preset feature of the second recall result sample based on the second object profile information sample and the second query information sample, thereby obtaining the feature scores for each preset feature of the second recall result sample. The ranking score label of the second recall result sample is determined based on the feature weights and feature scores of each preset feature of the second recall result sample.

9. The recall result ranking method as described in claim 8, characterized in that, The ranking and scoring labels of the second recall result samples are determined based on the feature weights and feature scores of each preset feature, including: The weight of each preset feature is multiplied by its corresponding feature score to obtain the weighted score of each preset feature. The weighted scores of the multiple preset features are summed to obtain the ranking score label of the second recall result.

10. A recall result sorting device, characterized in that, The device includes: The acquisition unit is used to acquire query information, object profile information of the query object corresponding to the query information, and multiple recall results corresponding to the query information; The annotation unit is used to annotate each of the multiple recall results based on the query information and the object profile information to obtain annotation information for each recall result. The annotation information is used to indicate the degree of matching between each recall result and the query information and / or the object profile information. The sorting unit is used to sort the multiple recall results based on the annotation information of each recall result, the query information, and the object profile information.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the recall result sorting method as described in any one of claims 1-9.

12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the recall result sorting method as described in any one of claims 1-9 by calling the computer program stored in the memory.

13. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the recall result sorting method according to any one of claims 1-9.