Content querying method and apparatus, device, storage medium, and program product

By using machine learning models to determine user query patterns and generate target content with specific visual styles, the problem of low browsing efficiency caused by excessive content on traditional query result pages is solved, thus improving information retrieval efficiency.

WO2026044466A1PCT designated stage Publication Date: 2026-03-05BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/114646
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

When traditional search results pages display a large amount of content, user browsing efficiency is poor, affecting the completeness of information retrieval.

Method used

The machine learning model determines the matching degree between the user query and multiple query patterns, generates and presents target content of a predetermined type, extracts target content that matches the user query from the content database, and presents it in the query results page with a specific visual style.

Benefits of technology

It improves the efficiency of users obtaining content, making it easier for them to find the answers they need on the search results page.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114646_05032026_PF_FP_ABST
    Figure CN2024114646_05032026_PF_FP_ABST
Patent Text Reader

Abstract

According to the embodiments of the present disclosure, provided are a content querying method and apparatus, a device, a storage medium, and a program product. The method comprises: in response to receiving a user query, determining multiple degrees of matching between the user query and multiple query patterns; based on the multiple degrees of matching, determining whether a query result for the user query should include content of a predetermined type, the content of the predetermined type being generated using a machine learning model based on at least one data source; in response to determining that the query result for the user query should include content of the predetermined type, extracting target content from a content database that matches the user query, the content database comprising content of the predetermined type; and displaying the target content on a query result page for the user query with a visual style corresponding to the predetermined type. Accordingly, the efficiency of obtaining content by users can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, equipment, storage media, and program products for content retrieval Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for content retrieval. Background Technology

[0002] With the development of information technology, various terminal devices can provide people with a variety of services in work and life. For example, terminal devices can be equipped with applications that provide services. Terminal devices or applications can provide users with content query functions, content browsing functions, etc., to assist users in using the terminal devices or applications. Applications can provide various types of pages, and through these pages, receive user queries and provide users with query result pages corresponding to the user queries.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for content retrieval is provided. The method includes: in response to receiving a user query, determining multiple matching degrees between the user query and multiple query patterns; based on the multiple matching degrees, determining whether query results for the user query should include content of a predetermined type, the predetermined type of content being generated using a machine learning model based on at least one data source; in response to determining that query results for the user query should include content of the predetermined type, extracting target content matching the user query from a content database, the content database including content of the predetermined type; and in a query results page for the user query, presenting the target content in a visual style corresponding to the predetermined type.

[0005] In a second aspect of this disclosure, an apparatus for content querying is provided. The apparatus includes: a matching degree determination module configured to, in response to receiving a user query, determine multiple matching degrees between the user query and multiple query patterns; a content determination module configured to, based on the multiple matching degrees, determine whether query results for the user query should include content of a predetermined type, the predetermined type of content being generated using a machine learning model based on at least one data source; a content extraction module configured to, in response to determining that query results for the user query should include content of the predetermined type, extract target content matching the user query from a content database, the content database including content of the predetermined type; and a content presentation module configured to, on a query results page for the user query, render the target content in a visual style corresponding to the predetermined type.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] Figure 2 illustrates a schematic diagram of an architecture for content retrieval according to some embodiments of the present disclosure;

[0013] Figure 3 shows a schematic diagram of an example query results page according to some embodiments of the present disclosure;

[0014] Figure 4 shows a flowchart of a content query method according to some embodiments of the present disclosure;

[0015] Figure 5 illustrates an exemplary structural block diagram of an apparatus for content retrieval according to some embodiments of the present disclosure; and

[0016] Figure 6 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0026] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0027] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as an input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. The testing phase can sometimes be integrated into the training phase. In the application or inference phase, the trained model can be used to process actual model inputs based on the trained parameter values ​​to determine the corresponding model output.

[0028] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 115 is installed on a client device 110. A user 140 can interact with the application 115 via the client device 110 and / or an attached device of the client device 110. For example, the application 115 can capture the user 140's voice via an audio capture device (e.g., a microphone) of the client device 110, can capture the user 140's image via an image capture device (e.g., a camera) of the client device 110, and so on.

[0029] In embodiments of this disclosure, application 115 can be any suitable application with search functionality (i.e., content query functionality), such as a browser application, social networking application, media item application, etc. In environment 100, if application 115 is active, client device 110 can display page 150 of application 115. Page 150 can include various types of pages that application 115 can provide, such as search pages, query results pages, content browsing pages, user profile pages, etc. In some embodiments, client device 110 and / or application 115 can receive user queries and provide query results corresponding to user queries via page 150.

[0030] In some embodiments, a communication connection is established between the client device 110 and the server device 120. The communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections, and the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, the client device 110 and the server device 120 can perform signaling interaction through their communication connection to provide services to application 115.

[0031] As shown in Figure 1, server device 120 can invoke machine learning model 130 to support the search function of application 115 based on the output of machine learning model 130. Machine learning model 130 can be deployed on server device 120 or on other devices. Machine learning model 130 can be based on any suitable model architecture, including but not limited to Transformer models, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), etc. In some embodiments, machine learning model 130 can be based on a language model (LM). Language models, by learning from a large corpus, are capable of question answering. Machine learning model 130 can also be based on other suitable models. It should be noted that machine learning model 130 may include one or more machine learning models. If machine learning model 130 includes multiple machine learning models, these multiple machine learning models may have different uses and functions, which is not limited in this disclosure.

[0032] Client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, client device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0033] The server-side device 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server-side device 120 may include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in a cloud environment, etc.

[0034] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0035] Traditionally, in response to a user query, the search results page displays search results that include multiple items. For example, if a user query specifies a media item, the search results page could display results for multiple media items. If the search results contain a large amount of content, the user's browsing efficiency may be poor, potentially affecting the completeness of the information obtained.

[0036] In view of this, according to embodiments of the present disclosure, an improved content query scheme is provided. According to the scheme of the embodiments of the present disclosure, in response to receiving a user query, multiple matching degrees between the user query and multiple query patterns are determined. Based on the multiple matching degrees, it is determined whether the query results for the user query should include content of a predetermined type, the predetermined type of content being generated using a machine learning model based on at least one data source. In response to determining that the query results for the user query should include content of the predetermined type, target content matching the user query is extracted from a content database, the content database including content of the predetermined type. In the query results page for the user query, the target content is presented in a visual style corresponding to the predetermined type.

[0037] In this way, target content matching the user's query can be extracted from the content database and presented in a specific visual style on the query results page. This allows users to easily obtain the answer (i.e., the target content) to their query on the query results page, improving the efficiency of content retrieval.

[0038] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0039] Figure 2 illustrates a schematic diagram of an architecture 200 for content querying according to some embodiments of the present disclosure. For ease of description, the architecture 200 is illustrated as being implemented at a server device 120. It should be noted that if the architecture 200 is illustrated as being implemented at a client device 110, some operations described with reference to the client device 110 may require the assistance of the server device 120. It should be noted that the operations performed by the client device 110 may specifically be performed by relevant applications installed on the client device 110. The architecture 200 will be described with reference to the environment 100 of Figure 1. The architecture 200 involves a matching degree determination unit 220, a determination unit 230, an extraction unit 240, and a presentation unit 260.

[0040] Client device 110 can receive user query 210 from user 140 via page 150. User query 210 can be any appropriate type of query, including but not limited to voice, action, image, and text queries. Client device 110 can provide the received user query 210 to server device 120 via communication with server device 120. Server device 120 can provide user query 210 to matching determination unit 220 in response to receiving user query 210.

[0041] The matching degree determination unit 220 can, in response to receiving a user query 210, determine multiple matching degrees between the user query 210 and multiple query patterns. The multiple query patterns can include at least a first query pattern 222 (also known as a resource-deficient pattern) indicating whether the query results can satisfy the user's needs corresponding to the query, a second query pattern 227 (also known as a strong question-answering pattern) indicating whether the user query is related to knowledge-based question answering, and a third query pattern 229 (also known as an information-finding pattern) indicating whether the user query is related to information retrieval. Of course, the multiple query patterns can also include any other suitable query patterns, and this disclosure does not limit this. It is generally found that presenting automatically generated content to the user in a special visual style under certain query patterns is more beneficial. Therefore, for the current user query, it can be determined whether automatically generated content needs to be provided on the user query results page by judging whether the user query matches certain predefined query patterns.

[0042] Regarding the specific method for determining multiple matching degrees corresponding to multiple query patterns, in some embodiments, the matching degree determination unit 220 can determine the first matching degree between user query 210 and the first query pattern 222 by using a trained first machine learning model 221 (e.g., a click prediction model) to determine the predicted probability 225 of the query result 223 being clicked on the query results page for user query 210. The predicted probability 225 and the first matching degree can be positively correlated, that is, the higher the predicted probability 225 of the query result 223 being clicked, the higher the first matching degree. This is because if it is predicted that the user is more willing to click on the matching search results on the search results page, then it may no longer be necessary to provide content automatically generated and summarized by the model on the search results page.

[0043] Specifically, the matching degree determination unit 220 can obtain multiple query results 223 that match the user query 210 and are presented on the query results page. Each query result 223 can include content of any appropriate type, such as documents, web pages, media items (e.g., images, videos, audio, etc.), etc. It should be noted that the multiple query results 223 presented on the query results page can be a subset of all query results that match the user query 210. For example, if all query results that match the user query 210 include 100 query results, the multiple query results presented on the query results page can be 20 of these 100 query results. These 20 query results can be 20 randomly selected from the 100 query results, or they can be the 20 query results with the highest matching degree to the user query 210 from the 100 query results. It can be understood that the number of multiple query results 223 can be associated with the configuration of the query results page. For example, if the query results page is configured to display 20 query results at a time, the number of multiple query results 223 can be 20. If the query results page is configured to display 50 query results at a time, the number of multiple query results 223 can be 50.

[0044] The matching degree determination unit 220 can extract at least one type of feature information 224 from each of the multiple query results 223. The at least one type of feature information 224 for each query result 223 may include, for example, the click-through rate (CTR), relevance, authority, etc. The matching degree determination unit 220 can provide the at least one type of feature information 224 from each of the multiple query results 223 to the first machine learning model 221. The first machine learning model 221 can, for example, determine the predicted probability 225 of the query result 223 being clicked on the query results page based on the at least one type of feature information 224 from each of the multiple query results 223. The output of the first machine learning model 221 can, for example, be a value between 0 and 1. The matching degree determination unit 220 can determine the predicted probability 225 of the query result 223 being clicked on the query results page based on this value (for example, if the model output is 0.7, then the predicted probability 225 is 70%), thereby determining the first matching degree between the user query 210 and the first query pattern 222.

[0045] In some embodiments, the matching degree determination unit 220 may use a trained second machine learning model 226 (e.g., a strong question-answering intent prediction model) to determine a second matching degree between user query 210 and second query pattern 227, and may use a trained third machine learning model 228 (e.g., an information knowledge intent model) to determine a third matching degree between user query 210 and third query pattern 229. For example, the second machine learning model 226 and the third machine learning model 228 may each output a value between 0 and 1 based on user query 210, and the matching degree determination unit 220 may determine the second and third matching degrees based on the values ​​output by the second machine learning model 226 and the third machine learning model 228, respectively. It should be noted that the first machine learning model 221, the second machine learning model 226, and the third machine learning model 228 can all be machine learning models included in machine learning model 130, and these three machine learning models can be based on any suitable model structure. This is because, under a strong question-and-answer intent or information-pointing diagram, users may be more inclined to search for knowledge or information, or to seek a certain fact. In this case, providing summary content automatically generated by the model and matching the user's query on the search results page may be more desirable for the user.

[0046] In some embodiments, the sample set used to train the second machine learning model 226 and the sample set used to train the third machine learning model 228 may each include multiple sample queries and multiple labels. Each label in the sample set used to train the second machine learning model 226 may indicate the label matching degree between the corresponding sample query and the second query pattern 227, and each label in the sample set used to train the third machine learning model 228 may indicate the label matching degree between the corresponding sample query and the third query pattern 229. It should be noted that these two sample sets can be the same, in which case each label in the sample set may indicate the label matching degree between the corresponding sample query and the second query pattern 227 and the third query pattern 229. These two sample sets can also be different, in which case the multiple sample queries in the two sample sets may be the same, partially the same, or different. Furthermore, even if the multiple sample queries in the two sample sets are the same, the labels corresponding to the same sample query will be different in the two sample sets.

[0047] For ease of description, the sample set used to train the second machine learning model 226 and the sample set used to train the third machine learning model 228 will be collectively referred to as the target sample set. Each label in the target sample set may be pre-determined by the user manually (e.g., manual labeling). In some embodiments, to save labor costs, the training devices used to train the second machine learning model 226 and / or the third machine learning model 228 (which may be, for example, server device 120 or any other suitable electronic device) may acquire a first set of sample queries and the label matching degree corresponding to the first set of sample queries. The number of sample queries included in the first set of sample queries is less than the number of sample queries included in the target sample set. For example, the first set of sample queries may include 2000 sample queries, and the target sample set may include 20000 sample queries.

[0048] The training device can provide a first set of sample queries and a first prompt word input to a trained language model. The language model, based on the first prompt word input, determines the expected matching degree between each of the first set of sample queries and the second and / or third query patterns. It can be understood that if the target sample set is the sample set used to train the second machine learning model 226, the first set of sample queries is also a set of sample queries used to train the second machine learning model 226, and the training device can use the language model to determine the expected matching degree between each of the first set of sample queries and the second query pattern. If the target sample set is the sample set used to train the third machine learning model 228, the first set of sample queries is also a set of sample queries used to train the third machine learning model 228, and the training device can use the language model to determine the expected matching degree between each of the first set of sample queries and the third query pattern. If the target sample set is both the sample set used to train the second machine learning model 226 and the sample set used to train the third machine learning model 228, then the first set of sample queries is also both the sample set used to train the second machine learning model 226 and the sample set used to train the third machine learning model 228. The training device can use the language model to determine the expected matching degree between the first set of sample queries and the second and third query patterns, respectively.

[0049] The first prompt word input indicates the scoring requirement of the language model for the first set of sample queries. It can be understood that the first prompt word input used to determine the expected match between each of the first set of sample queries and the second query pattern (referred to as the first prompt word input corresponding to the second query pattern) and the first prompt word input used to determine the expected match between each of the first set of sample queries and the third query pattern (referred to as the first prompt word input corresponding to the third query pattern) are different. The first prompt word input corresponding to each query pattern can indicate how to determine the score for the first set of sample queries for that query pattern.

[0050] The training device can determine the difference between the predicted matching degree and the labeled matching degree of each of the first set of sample queries output by the language model, and adjust the first prompt word input based on this difference. For example, both the predicted matching degree and the labeled matching degree can be values ​​between 0 and 1. The training device can determine the difference between the value corresponding to the predicted matching degree and the value corresponding to the labeled matching degree; this difference is the discrepancy between the predicted and labeled matching degrees. The adjustment goal of the first prompt word input is to reduce the discrepancy, i.e., to decrease the difference. The training device can determine that the adjustment of the first prompt word is complete when the difference is less than a threshold (e.g., 0).

[0051] The training device can use a language model, based on the adjusted first prompt word input, to determine the expected matching degree between each of the second set of sample queries and the second and / or third query patterns, as the labeled matching degree of the second set of sample queries. The second set of sample queries can be obtained in any suitable manner. In some embodiments, the training device can obtain the second set of sample queries using a language model and the first set of sample queries. The second set of sample queries can be sample queries generated by the language model with reference to the first set of sample queries. The number of sample queries included in the second set of sample queries can, for example, be greater than the number of sample queries included in the first set of sample queries. The process by which the training device generates the second set of sample queries based on the first set of sample queries can also be referred to as an extension of the first set of sample queries.

[0052] The training device can provide the adjusted first prompt word input and the second set of sample queries to the language model, and determine the expected matching degree between each of the second set of sample queries and the second and / or third query patterns based on the output of the language model. The training device can then define the expected matching degree of each of the second set of sample queries as its corresponding labeled matching degree. For example, the training device can use the first and second sets of sample queries to determine the target sample set. That is, multiple sample queries in the target sample set represent the total number of the first and second sets of sample queries. The label corresponding to each sample query in the target sample set can be, for example, the labeled matching degree of that sample query.

[0053] Matching degree determination unit 220 can provide multiple determined matching degrees to determination unit 230. Determination unit 230 can determine, based on the multiple matching degrees, whether the query results for user query 210 should include content of a predetermined type. The predetermined type of content may be generated using a machine learning model (e.g., machine learning model 130) based on at least one data source. At least one data source may include, but is not limited to, web pages, images, documents, videos, etc. In some embodiments, determination unit 230 can obtain a first matching degree threshold 231 corresponding to each matching degree, and determine whether the query results should include content of a predetermined type based on the comparison result of each matching degree among the multiple matching degrees with the corresponding first matching degree threshold 231.

[0054] In some embodiments, the determining unit 230 may determine that the query result should include content of a predetermined type in response to determining that at least one of the multiple matching degrees satisfies its corresponding first matching degree threshold 231. The case where the matching degree corresponding to a certain query pattern satisfies its corresponding first matching degree threshold can be referred to as user query 210 matching that query pattern. User query 210 may match at least one query pattern. For example, if the multiple query patterns include three query patterns, the matching degree determining unit 220 may determine the three matching degrees corresponding to these three query patterns. The determining unit 230 may determine the three first matching degree thresholds 231 corresponding to these three query patterns, and may determine that the query result should include content of a predetermined type in response to any one of the three query patterns satisfying its corresponding first matching degree threshold 231. Of course, if any two query patterns satisfy their corresponding first matching degree thresholds 231, or if all three query patterns satisfy their corresponding first matching degree thresholds 231, the determining unit 230 may still determine that the query result should include content of a predetermined type. For ease of description, the following example illustrates the scenario where user query 210 only has one matching degree among multiple matching degrees for multiple query patterns that satisfies the corresponding first matching degree threshold (i.e., user query 210 only matches one of the multiple query patterns).

[0055] It should be noted that, in some embodiments, for the first query mode 222, the determining unit 230 may determine that the first matching degree of the first query mode 222 meets the first matching degree threshold 231 corresponding to the first query mode 222 in response to the first matching degree of the first query mode 222 not reaching the first matching degree threshold 231 corresponding to the first query mode 222. That is, the determining unit 230 may determine that the user query 210 matches the first query mode 222 even when the first matching degree is relatively low.

[0056] For the second query mode 227 and / or the third query mode 229, the determining unit 230 can determine that the second matching degree of the second query mode 227 satisfies the first matching degree threshold 231 corresponding to the second query mode 227, and / or the third matching degree of the third query mode 229 satisfies the first matching degree threshold 231 corresponding to the third query mode 229, in response to the second matching degree of the second query mode 227 reaching the first matching degree threshold 231 corresponding to the second query mode 227, and / or the third matching degree of the third query mode 229 satisfying the first matching degree threshold 231 corresponding to the third query mode 229. That is, the determining unit 230 can determine that the user query 210 matches the second query mode 227 and / or the third query mode 229 when the second matching degree and / or the third matching degree are relatively large.

[0057] Extraction unit 240 may, in response to determining that the query results for user query 210 should include content of a predetermined type, extract target content matching user query 210 from content database 245, which includes content of the predetermined type. The content in content database 245 may be pre-generated by server device 120 using a machine learning model (e.g., machine learning model 130). Alternatively or additionally, in some embodiments, the content in content database 245 may also be pre-generated by other electronic devices using machine learning models. In this case, server device 120 may obtain content database 245, which includes content of the predetermined type, from other electronic devices via a communication connection. For ease of description, the following example illustrates that the content in content database 245 is generated by server device 120 using a machine learning model.

[0058] Regarding the generation method of the content in content database 245, in some embodiments, server device 120 may utilize a machine learning model (e.g., machine learning model 130) to generate a first answer and a second answer matching the reference query. The content included in the first answer may be more detailed than that included in the second answer. The first answer may be referred to as a long answer, and the second answer may be referred to as a short answer. Server device 120 may determine the quality scores of the first answer and the second answer respectively, and based on the quality scores of the first answer and the second answer, determine a retention strategy for the first answer and the second answer. Server device 120 may use any suitable method to determine the quality scores of the first answer and the second answer respectively, and this disclosure does not limit the specific method of determining the quality scores. The retention strategy may indicate whether to retain the corresponding answer.

[0059] In some embodiments, server device 120 can determine the quality scores of the first and second answers by detecting whether predetermined search terms are included in the first and second answers. Predetermined search terms may be, for example, search terms that indicate the machine learning model cannot generate an accurate answer for user query 210. For example, predetermined search terms may include, but are not limited to, "sorry," "apologies," "unable to generate," "unable to find," etc. Server device 120 may determine that the quality score of the first / second answer is poor in response to the inclusion of predetermined search terms in the first / second answer, and may determine that the quality score of the first / second answer is high in response to the absence of predetermined search terms in the first / second answer. For example, server device 120 may determine that the retention policy corresponding to a poor quality score for either the first or second answer indicates that the answer should not be retained.

[0060] Server-side device 120 may, for example, save the first and second answers to content database 245 as content of a predetermined type matching the reference query, only if a retention policy indicates that both the first and second answers are to be retained. Server-side device 120 may, for example, save the reference query, the first answer, and the second answer in the content database 245 in the form of "reference query - first answer - second answer". In some embodiments, server-side device 120 may also determine the semantic differences between the first and second answers corresponding to each reference query. Server-side device 120 may retain both answers only if the semantic difference between the first and second answers is less than a threshold. Thus, server-side device 120 can perform cross-validation on the first and second answers corresponding to the reference query, which can improve the accuracy of the first and second answers.

[0061] Regarding the specific method of extracting target content matching user query 210 from content database 245, in some embodiments, server device 120 may perform a search in content database 245 based on user query 210 to find reference queries similar to user query 210. For example, server device 120 may determine the similarity between multiple reference queries in content database 245 and user query 210. For instance, server device 120 may determine the content corresponding to a group of reference queries (which may include at least one reference query) with a similarity higher than a threshold (e.g., the first and second answers corresponding to each query in this group) as target content matching user query 210.

[0062] In some embodiments, the server device 120 can directly provide the target content to the presentation unit 260. The presentation unit 260 can present the target content in a visual style corresponding to a predetermined type on the query results page for the user query. The visual style corresponding to the predetermined type can at least include a card style. That is, the presentation unit 260 can present the target content in card form on the query results page. In some embodiments, regarding the presentation position of the target content on the query results page, the server device 120 can determine the presentation position of the target content on the query results page based at least on the user query 210 and multiple matching degrees corresponding to multiple query patterns.

[0063] Each query pattern can correspond to multiple matching thresholds. In some embodiments, the first matching threshold 231 corresponding to the first query pattern 222 can be the largest matching threshold among the multiple matching thresholds corresponding to the first query pattern 222. Taking the first query pattern 222 as an example, which can correspond to three matching thresholds (e.g., the first matching threshold 231, the second matching threshold, and the third matching threshold), the order of these three matching thresholds can be: first matching threshold 231 > second matching threshold > third matching threshold. The server device 120 (specifically, for example, the determining unit 230) can determine that the query result of the user query 210 should include content of a predetermined type in response to the first matching degree corresponding to the first query pattern 222 satisfying the first matching threshold 231 (i.e., not reaching the first matching threshold 231).

[0064] For the first query pattern 222, a smaller matching degree threshold corresponds to a higher priority position on the query results page. Taking the first position as the top position on the query results page, the second position as the position below the first position, and the third position as the position below the second position as an example, if the first matching degree corresponding to the first query pattern 222 only meets the corresponding first matching degree threshold 231 (e.g., less than the first matching degree threshold 231 but greater than the second matching degree threshold), the presentation unit 260 can determine that the target content is presented in the third position on the query results page. If the matching degree corresponding to the query pattern meets the second matching degree threshold but does not meet the third matching degree threshold (e.g., less than the second matching degree threshold but greater than the third matching degree threshold), the presentation unit 260 can determine that the target content is presented in the second position on the query results page. If the matching degree corresponding to the query pattern meets the third matching degree threshold (e.g., less than the third matching degree threshold), the presentation unit 260 can determine that the target content is presented in the first position on the query results page.

[0065] Similarly, for any of the query modes 227 and 229, the first matching threshold 231 can be, for example, the smallest matching threshold among multiple matching thresholds corresponding to that query mode. Taking the query mode as an example where it corresponds to three matching thresholds (e.g., the first matching threshold 231, the second matching threshold, and the third matching threshold), the order of these three matching thresholds can be: first matching threshold 231 < second matching threshold < third matching threshold. The server device 120 (specifically, for example, the determining unit 230) can determine that the query result of user query 210 should include content of a predetermined type in response to the matching degree corresponding to the query mode satisfying the first matching threshold 231 (i.e., reaching the first matching threshold 231).

[0066] For either the second query mode 227 or the third query mode 229, a higher matching degree threshold corresponds to a higher priority position on the query results page. Taking the first position as the topmost position on the query results page, the second position as the position below the first position, and the third position as the position below the second position as an example: if the matching degree corresponding to the query mode only meets the corresponding first matching degree threshold 231 (e.g., greater than the first matching degree threshold 231 but less than the second matching degree threshold), the presentation unit 260 can determine that the target content is presented in the third position on the query results page. If the matching degree corresponding to the query mode meets the second matching degree threshold but does not meet the third matching degree threshold (e.g., greater than the second matching degree threshold but less than the third matching degree threshold), the presentation unit 260 can determine that the target content is presented in the second position on the query results page. If the matching degree corresponding to the query mode meets the third matching degree threshold (e.g., greater than the third matching degree threshold), the presentation unit 260 can determine that the target content is presented in the first position on the query results page.

[0067] In some embodiments, the presentation unit 260 may also determine the presentation position of the target content on the query results page based on the specific content of the target content (e.g., the text included in the target content), the query results corresponding to the user query 210, and the user's pre-configured information for the query results page. For example, the presentation unit 260 may determine whether the query results matching the user query 210 include content configured to be in a specified presentation position (e.g., user information, physical cards, featured answers, etc., corresponding to the user query 210). If it is determined that the query results matching the user query 210 include content configured to be in a specified presentation position, the presentation unit 260 may determine the presentation position of the target content to be in a position other than the specified presentation position based on at least multiple matching degrees. For example, if the query results include content configured to be in the first position, even if the presentation unit 260 determines the presentation position of the target content in the query results page to be in the first position based on the matching degree, the presentation unit 260, in order to avoid the content configured to be in the first position, will determine the presentation position of the target content to be in the second position below the first position or any other appropriate position.

[0068] In some embodiments, if it is determined by any other suitable means that additional content 251 also needs to be presented using a visual style (e.g., card style) corresponding to a predetermined type, then the architecture 200 may also involve a sorting unit 250. The additional content 251 may include, for example, at least one type of content, such as a single-document summary 252, a multi-document summary 253, multimedia content 254, etc. The sorting unit 250 may sort the target content and other content 251 based on multiple features corresponding to each of the target content and other content 251. These multiple features may include, for example, query features (e.g., intent, timeliness, authority), document features (e.g., usefulness, authenticity, readability), query-doc features (e.g., relevance, CTR), etc.

[0069] The sorting unit 250 can, for example, calculate the quality score corresponding to each piece of content in the target content and other content 251 using a predetermined fusion formula. The fusion formula could be, for example, Quality Score = (Parameter 1 + Feature 1). 参数2 +(parameter 1 + feature 2) 参数2 +(Parameter 1 + Feature 3) 参数2 +……+(parameter 1 + feature N) 参数2Here, parameters 1 and 2 can be pre-set parameters, and features 1 to N are multiple features required for sorting. The sorting unit 250 can then sort the target content and other content 251 based on the quality scores corresponding to each content. In some embodiments, the sorting unit 250 can determine the number of contents allowed to be presented in a predetermined visual style (e.g., card style) on the query results page (which can also be understood as determining how many cards are allowed to be presented on the query results page). This number can be pre-configured by the user or be a default value.

[0070] If only one card is allowed to be displayed on the query results page, the sorting unit can determine the content with the highest quality score from the target content and other content 251, and provide that content to the presentation unit 260 so that the content is presented on the query results page in a visual style corresponding to a predetermined type. Similarly, if multiple cards are allowed to be displayed on the query results page, the sorting unit can determine multiple pieces of content with the highest quality scores from the target content and other content 251 based on the number of cards, and provide these multiple pieces of content to the presentation unit 260 so that these multiple pieces of content are presented on the query results page in a visual style corresponding to a predetermined type.

[0071] Figure 3 illustrates a schematic diagram of an example 300 of a query results page according to some embodiments of the present disclosure. As shown in Figure 3, example 300 includes an input box 310, a card 320, and a region 330. Client device 110 may receive a user query, for example, via the input box 310. Client device 110 may present the query results corresponding to the user query, for example, in region 330. Card 320 may display target content extracted from a content database by server device 120. If the required display size of the target content exceeds the size of card 320, card 320 may include a control 321. Client device 110 may, in response to receiving a trigger operation on control 321, switch to a details page of the target content and present the target content in the details page. Card 320 may also include a region 322. Region 322 may display at least one data source of the target content (e.g., data source 1 to data source 4 shown in the figure). If the number of at least one data source is large, region 322 may also display only a portion of the data sources from at least one data source. In this case, region 322 may include a control 323. Client device 110 can respond to a received trigger operation on control 323 by presenting the entirety of at least one data source.

[0072] In summary, according to the embodiments of this disclosure, target content matching a user's query can be extracted from a content database and presented in a specific visual style on the query results page. This allows users to conveniently obtain the answer (i.e., the target content) corresponding to their query on the query results page, improving the efficiency of content retrieval for users.

[0073] Figure 4 shows a flowchart of a content query method 400 according to some embodiments of the present disclosure. Method 400 may be implemented at a server device 120. Method 400 will be described with reference to the environment 100 of Figure 1.

[0074] In box 410, server device 120 responds to receiving a user query by determining multiple matching degrees between the user query and multiple query patterns.

[0075] In box 420, server device 120 determines, based on multiple matching degrees, whether the query results for a user query should include content of a predetermined type, which is generated using a machine learning model based on at least one data source.

[0076] In box 430, in response to determining that the query results for the user query should include content of a predetermined type, the server device 120 extracts target content that matches the user query from a content database, which includes content of the predetermined type.

[0077] In box 440, server device 120 renders the target content in a visual style corresponding to a predetermined type on the query results page for the user query.

[0078] In some embodiments, determining multiple matching degrees between a user query and multiple query patterns includes: determining a first matching degree between the user query and a first query pattern by using a trained first machine learning model to determine the expected probability that a query result in a query results page for the user query will be clicked, the first query pattern indicating whether the query result can meet the user's needs corresponding to the user query; determining a second matching degree between the user query and a second query pattern by using a trained second machine learning model, the second query pattern indicating that the user query is related to knowledge question answering; and determining a third matching degree between the user query and a third query pattern by using a trained third machine learning model, the third query pattern indicating that the user query is related to information retrieval.

[0079] In some embodiments, determining a first matching degree between a user query and a first query pattern includes: obtaining multiple query results that match the user query, the multiple search results being presented on a query results page; extracting at least one type of feature information from each of the multiple query results; and using a first machine learning model, based on at least one type of feature information from each of the multiple query results, determining the expected probability that the query results on the query results page will be clicked.

[0080] In some embodiments, the second machine learning model and / or the third machine learning model are trained based on a target sample set, which includes multiple sample queries and multiple labels, each label indicating the label matching degree between the corresponding sample query and the second query pattern, and / or the label matching degree between the corresponding sample query and the third query pattern.

[0081] In some embodiments, the target sample set is obtained by: using a language model to determine the expected matching degree between each of the first set of sample queries and the second query pattern and / or the third query pattern based on the first prompt word input, wherein the first prompt word input indicates the scoring requirement of the language model for the first set of sample queries, and the first set of sample queries has a corresponding labeled matching degree; adjusting the first prompt word input based on the difference between the labeled matching degree and the expected matching degree corresponding to each of the first set of sample queries; and using the language model to determine the expected matching degree between each of the second set of sample queries and the second query pattern and / or the third query pattern based on the adjusted first prompt word input, as the labeled matching degree of the second set of sample queries.

[0082] In some embodiments, determining whether the query results for a user query should include content of a predetermined type based on multiple matching degrees includes: in response to determining that at least one of the multiple matching degrees satisfies its corresponding first matching degree threshold, determining that the query results should include content of a predetermined type.

[0083] In some embodiments, presenting target content in a visual style corresponding to a predetermined type includes: determining the presentation position of the target content on a query results page based on at least a plurality of matching degrees; and presenting the target content in a visual style corresponding to a predetermined type at the presentation position on the query results page.

[0084] In some embodiments, determining the presentation position of target content on the query results page includes: determining whether the query results matching the user query include content configured to be in a specified presentation position; and if the query results matching the user query include content configured to be in a specified presentation position, determining the presentation position of the target content to be in a different presentation position than the specified presentation position based on at least multiple matching degrees.

[0085] In some embodiments, content of a predetermined type in the content database is generated by: generating a first answer and a second answer that match a reference query using a machine learning model, wherein the first answer contains more detailed content than the second answer; determining a retention strategy for the first answer and the second answer based on their respective quality scores; and saving the first answer and the second answer to the content database as content of a predetermined type that matches the reference query if the retention strategy indicates that both the first answer and the second answer should be retained.

[0086] In some embodiments, the visual style corresponding to the predetermined type includes at least a card style.

[0087] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows an exemplary structural block diagram of an apparatus 500 for content retrieval according to some embodiments of this disclosure. The apparatus 500 may be implemented as or included in the server device 120. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0088] As shown in Figure 5, the device 500 includes a matching degree determination module 510, configured to determine multiple matching degrees between the user query and multiple query patterns in response to receiving a user query. The device 500 also includes a content determination module 520, configured to determine, based on the multiple matching degrees, whether the query results for the user query should include content of a predetermined type, the predetermined type of content being generated using a machine learning model based on at least one data source. The device 500 also includes a content extraction module 530, configured to extract target content matching the user query from a content database, the content database including content of the predetermined type, in response to determining that the query results for the user query should include content of the predetermined type. The device 500 also includes a content presentation module 540, configured to present the target content in a visual style corresponding to the predetermined type on the query results page for the user query.

[0089] In some embodiments, the matching degree determination module 510 is further configured to: determine a first matching degree between the user query and a first query pattern by using a trained first machine learning model to determine the expected probability of a query result being clicked on a query results page for a user query, the first query pattern indicating whether the query result can meet the user's needs corresponding to the user query; determine a second matching degree between the user query and a second query pattern by using a trained second machine learning model, the second query pattern indicating that the user query is related to knowledge question answering; and determine a third matching degree between the user query and a third query pattern by using a trained third machine learning model, the third query pattern indicating that the user query is related to information retrieval.

[0090] In some embodiments, the matching degree determination module 510 is further configured to: obtain multiple query results that match the user query, the multiple search results to be presented on the query results page; extract at least one type of feature information for each of the multiple query results; and use a first machine learning model to determine the expected probability that the query results on the query results page will be clicked based on at least one type of feature information for each of the multiple query results.

[0091] In some embodiments, the second machine learning model and / or the third machine learning model are trained based on a target sample set, which includes multiple sample queries and multiple labels, each label indicating the label matching degree between the corresponding sample query and the second query pattern, and / or the label matching degree between the corresponding sample query and the third query pattern.

[0092] In some embodiments, the target sample set is obtained by: using a language model to determine the expected matching degree between each of the first set of sample queries and the second query pattern and / or the third query pattern based on the first prompt word input, wherein the first prompt word input indicates the scoring requirement of the language model for the first set of sample queries, and the first set of sample queries has a corresponding labeled matching degree; adjusting the first prompt word input based on the difference between the labeled matching degree and the expected matching degree corresponding to each of the first set of sample queries; and using the language model to determine the expected matching degree between each of the second set of sample queries and the second query pattern and / or the third query pattern based on the adjusted first prompt word input, as the labeled matching degree of the second set of sample queries.

[0093] In some embodiments, the content determination module 520 is further configured to: determine that the query result should include content of a predetermined type in response to determining that at least one of the plurality of matching degrees meets its respective first matching degree threshold.

[0094] In some embodiments, the content presentation module 540 is further configured to: determine the presentation position of the target content on the query results page based on at least a plurality of matching degrees; and at the presentation position on the query results page, present the target content in a visual style corresponding to a predetermined type.

[0095] In some embodiments, the content presentation module 540 is further configured to: determine whether the query results matching the user query include content configured to be in a specified presentation position; and if it is determined that the query results matching the user query include content configured to be in a specified presentation position, determine the presentation position of the target content as another presentation position outside the specified presentation position based on at least multiple matching degrees.

[0096] In some embodiments, content of a predetermined type in the content database is generated by: generating a first answer and a second answer that match a reference query using a machine learning model, wherein the first answer contains more detailed content than the second answer; determining a retention strategy for the first answer and the second answer based on their respective quality scores; and saving the first answer and the second answer to the content database as content of a predetermined type that matches the reference query if the retention strategy indicates that both the first answer and the second answer should be retained.

[0097] In some embodiments, the visual style corresponding to the predetermined type includes at least a card style.

[0098] The units and / or modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0099] It should be understood that one or more steps in the above methods can be performed by appropriate electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may, for example, include the server device 120 in Figure 1.

[0100] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the server device 120 of Figure 1, or the apparatus 500 of Figure 5.

[0101] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to the program stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0102] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.

[0103] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0104] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0105] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0106] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0107] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0108] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0109] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some, as newer, implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0111] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for content retrieval, comprising: In response to receiving a user query, determine multiple matching degrees between the user query and multiple query patterns; Based on the multiple matching degrees, it is determined whether the query results for the user query should include content of a predetermined type, which is generated using a machine learning model based on at least one data source; In response to determining that the query results for the user query should include content of the predetermined type, target content matching the user query is extracted from a content database, the content database including content of the predetermined type; as well as In the query results page for the user query, the target content is presented in a visual style corresponding to the predetermined type.

2. The method according to claim 1, wherein determining the multiple matching degrees between the user query and multiple query patterns includes: By using a trained first machine learning model to determine the expected probability that a query result in the query results page for the user query will be clicked, a first matching degree between the user query and a first query pattern is determined, wherein the first query pattern indicates whether the query results can meet the user's needs corresponding to the user query. Using a trained second machine learning model, a second matching degree is determined between the user query and a second query pattern, the second query pattern indicating that the user query is related to knowledge question answering; as well as Using a trained third machine learning model, a third matching degree is determined between the user query and a third query pattern, which indicates that the user query is relevant to information retrieval.

3. The method according to claim 2, wherein determining the first matching degree between the user query and the first query pattern includes: Obtain multiple query results that match the user's query, and the multiple search results shall be presented on the query results page; Extract at least one type of feature information from each of the multiple query results; as well as Using the first machine learning model, based on at least one type of feature information of each of the multiple query results, the expected probability of a query result being clicked on the query results page is determined.

4. The method according to claim 2, wherein the second machine learning model and / or the third machine learning model are trained based on a target sample set, the target sample set including multiple sample queries and multiple labels, each label indicating the label matching degree between the corresponding sample query and the second query pattern, and / or the label matching degree between the corresponding sample query and the third query pattern.

5. The method of claim 4, wherein the target sample set is obtained via: Using a language model, the expected matching degree between each of the first group of sample queries and the second query pattern and / or the third query pattern is determined based on the input of the first prompt word. The input of the first prompt word indicates the scoring requirement of the language model for the first group of sample queries, and the first group of sample queries has a corresponding labeled matching degree. Based on the difference between the labeled matching degree and the expected matching degree of each of the first group of sample queries, the input of the first prompt word is adjusted; as well as Using the language model, based on the adjusted first prompt word input, the expected matching degree between each of the second group of sample queries and the second query pattern and / or the third query pattern is determined, and used as the labeled matching degree of the second group of sample queries.

6. The method of claim 1, wherein determining whether the query results for the user query should include content of a predetermined type based on the plurality of matching degrees includes: In response to determining that at least one of the multiple matching degrees satisfies its respective first matching degree threshold, it is determined that the query result should include content of the predetermined type.

7. The method of claim 1, wherein rendering the target content in a visual style corresponding to the predetermined type comprises: The position of the target content on the query results page is determined based on at least the multiple matching degrees. as well as At the specified display location on the query results page, the target content is aligned with the... The visual style corresponding to the predefined type is presented.

8. The method according to claim 7, wherein determining the presentation position of the target content on the query results page comprises: Determine whether the query results matching the user query include content configured to be displayed in a specified position; as well as If it is determined that the query results matching the user query include content configured to be in a specified presentation position, the presentation position of the target content is determined to be another presentation position other than the specified presentation position, based at least on the plurality of matching degrees.

9. The method of claim 1, wherein the content of the predetermined type in the content database is generated via: A machine learning model is used to generate a first and second answer that match a reference query, where the first answer contains more detailed information than the second answer. Based on the quality scores of the first answer and the second answer, a retention strategy is determined for the first answer and the second answer respectively; as well as If the retention policy indicates that both the first answer and the second answer are retained, the first answer and the second answer are saved to the content database as content of the predetermined type that matches the reference query.

10. The method of claim 1, wherein the visual style corresponding to the predetermined type includes at least a card style.

11. An apparatus for content retrieval, comprising: The matching degree determination module is configured to determine multiple matching degrees between the user query and multiple query patterns in response to receiving a user query; The content determination module is configured to determine, based on the multiple matching degrees, whether the query results for the user query should include content of a predetermined type, wherein the content of the predetermined type is generated using a machine learning model based on at least one data source; The content extraction module is configured to extract target content matching the user query from a content database, the content database including the content of the predetermined type, in response to determining that the query results for the user query should include content of the predetermined type. as well as The content presentation module is configured to appear on the query results page for the user's query. The target content is presented in a visual style corresponding to the predetermined type.

12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Information search method, apparatus and device, and storage medium

    CN108959531A

  • Query category speculation method and device, equipment and storage medium

    CN109753556A

  • Query request response method and device

    CN115114424A

  • Query method and device, equipment and storage medium

    CN116431039A

  • Generating a semantic search engine results page

    US20240256622A1