Search ranking method, device and question-answering system
By judging the timeliness of user query texts and sorting the search results based on their relevance and timeliness scores, the problem of poor accuracy of reply texts generated by large models is solved, and the accuracy of reply texts generated by large models is improved.
Patent Information
- Application Number
- CN202510600238.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the large model scenario, since the search engine does not care about the correctness of the search results, the accuracy of the reply text generated by the large model is poor, and the existing technology cannot effectively improve the accuracy of the reply text generated by the large model.
By determining whether the user query text is a time-sensitive query text, the search results are sorted based on the relevance score and timeliness score of the search results, and the sorting results are generated to assist the large model in generating reply text.
The accuracy of response text generated by large models has been improved, especially in the case of time-sensitive query text. By considering the timeliness score, the timeliness and relevance of the search results are combined; in the case of non-time-sensitive query text, only the relevance score is considered, reducing the adverse impact of the timeliness score on the ranking results and improving accuracy.
Smart Images

Figure CN120123500B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of search technology, and in particular to a search ranking method, device, and question-answering system. Background Art
[0002] Searching in a large model scenario is different from searching in traditional scenarios. In traditional scenarios, users judge the accuracy of search results and refer to the content themselves. In a large model scenario, the large model will understand the search results and give a response text. At the same time, due to the existence of large model hallucinations, the accuracy requirements for search results in large model scenarios will be higher.
[0003] Existing large-model search solutions mainly send the search engine's search results directly to the large model, and provide a reply text to the user's question based on the large model's understanding ability and the search results. The search effect is completely dependent on the search engine's search effect, and the search engine does not care about the correctness of the search results, resulting in poor accuracy of the reply text generated by the large model. Summary of the Invention
[0004] This application provides a search ranking method, device, and question-answering system to improve the accuracy of response text generated by large models.
[0005] According to a first aspect of an embodiment of the present application, a search ranking method is provided, comprising:
[0006] Obtaining search results based on the user's query text;
[0007] Determining whether the user query text is a time-sensitive query text;
[0008] In the case where the user query text is a time-sensitive query text, sorting the search results based on the relevance score and time-sensitive score of each search result to obtain a sorting result;
[0009] In the case where the user query text is not a time-sensitive query text, sorting the search results based on the relevance scores of the search results to obtain a sorting result;
[0010] The ranking result is used to assist the large model in generating a reply text corresponding to the user query text.
[0011] According to a second aspect of an embodiment of the present application, there is provided an electronic device, including a memory and a processor;
[0012] The memory is connected to the processor and is used to store programs;
[0013] The processor is configured to implement the search and sorting method as described in the first aspect by running the program in the memory.
[0014] According to a third aspect of an embodiment of the present application, a question-answering system is provided, including:
[0015] large models, search engines, and electronic devices;
[0016] The large model is configured to receive a user query text, and when determining that the user query text needs to be replied to through an information query, send the user query text to the search engine, send each search result fed back by the search engine to the electronic device, and generate a reply text corresponding to the user query text based on the feedback from the electronic device;
[0017] The search engine is configured to, upon receiving a user query text sent by the large model, perform a search based on the user query text, obtain search results, and feed the search results back to the large model;
[0018] The electronic device is used to process the search results sent by the large model according to the search ranking method as described in the first aspect when receiving the search results, and feed back the processing results to the large model.
[0019] In this application, the various search results obtained by searching based on the user query text are obtained, and it is first determined whether the user query text is a time-sensitive query text. If the user query text is a time-sensitive query text, it indicates that the timeliness of the search results can affect the accuracy of the search results. In the process of sorting the search results, the relevance score and timeliness score of the search results are comprehensively considered. Under the premise of ensuring the relevance sorting, the search results with strong timeliness are further given priority, which can improve the accuracy of the sorting results. The sorting results are used to assist the large model in generating the reply text corresponding to the user query text, further improving the accuracy of the reply text generated by the large model for the user query text. In the case that the user query text is not a time-sensitive query text, it indicates that the timeliness of the search results does not affect the accuracy of the search results. In the process of sorting the search results, the relevance score of the search results is considered, and the timeliness score of the search results is not considered. This reduces the adverse effect of the timeliness score on the sorting results, can improve the accuracy of the sorting results, and further improves the accuracy of the reply text generated by the large model for the user query text. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0021] Figure 1 A flowchart of a search and sorting method provided in an embodiment of the present application is shown.
[0022] Figure 2 This is a flow chart of step 102 provided in an embodiment of the present application.
[0023] Figure 3 This is a flow chart of step 201 provided in an embodiment of the present application.
[0024] Figure 4 This is a flow chart of step 202 provided in an embodiment of the present application.
[0025] Figure 5 This is a flow chart of step 203 provided in an embodiment of the present application.
[0026] Figure 6 This is a flow chart of step 103 provided in an embodiment of the present application.
[0027] Figure 7 This is a flow chart of step 104 provided in an embodiment of the present application.
[0028] Figure 8 This is a flowchart of step 701 provided in an embodiment of the present application.
[0029] Figure 9 A flowchart of a search and sorting method provided in an embodiment of the present application is shown.
[0030] Figure 10 This is a structural diagram of a search and sorting device provided in an embodiment of the present application.
[0031] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0032] Figure 12 This is a structural diagram of a question-answering system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] With the rise of big model technology, various AI companies have launched their own big model capability platforms. Users are increasingly relying on big model technology to help them acquire knowledge, assist with office work, generate content, and improve work efficiency. However, big model knowledge is integrated through large amounts of corpus training. This knowledge is limited and cannot keep up with rapidly changing information. Therefore, it is necessary to integrate third-party search sources to query content related to user questions on the public network to provide information support for the big model and help it answer user questions more accurately.
[0034] Searching in large-scale models differs from traditional search scenarios. In traditional search scenarios, users tend to enter relevant search keywords, while in large-scale models, users' questions are more likely to be expressed in a colloquial manner. In traditional search scenarios, users independently judge the accuracy of search results and refer to the content. In large-scale models, however, the large model understands the search results and provides a textual response. Furthermore, due to the phenomenon of large-scale model hallucinations, the accuracy of search results in large-scale models is even more demanding.
[0035] Existing large-scale search solutions primarily integrate with the service interfaces of relevant third-party search engines, breaking down user questions into search terms and inputting them into the service interfaces of the third-party search engines. The search results returned by the search engine are then sent directly to the large-scale model. Based on the large-scale model's understanding and search results, a textual response to the user's question is generated. The search engine's search results are primarily based on user click preferences or relevance, and the accuracy of the search results is not considered. Users independently refer to the accuracy of the search results and make their choices. However, in large-scale model scenarios, the large-scale model generally directly references the search results and generates a textual response. The search results are required to be relevant and accurate to avoid incorrect responses. Because the search engine does not care about the accuracy of the search results, the accuracy of the response text generated by the large-scale model is poor.
[0036] In order to improve the accuracy of reply text generated by large models, this application provides a search ranking method, device and question-answering system.
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] Exemplary Implementation Environment
[0039] The search ranking method according to the embodiments of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user device, a mobile device, a computing device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling computer-readable program instructions stored in a memory.
[0040] Exemplary Methods
[0041] See also Figure 1 In an exemplary embodiment, a search ranking method is provided. Figure 1 As shown, the process of the search ranking method mainly includes:
[0042] Step 101: Obtain search results based on a user query text.
[0043] In an exemplary embodiment, before step 101, the big model may first receive the user query text, and when it is determined that the user query text needs to be replied to through information query, the user query text is sent to the search engine, and the search engine searches based on the user query text, obtains each search result, and feeds back each search result to the big model. The big model then sends each search result fed back by the search engine to the electronic device, and the electronic device then executes steps 101 to 104.
[0044] For example, after the big model receives the user query text, it first uses its understanding ability to determine whether it needs to access the search engine for information query. It will only access the search engine when necessary. Other user query texts with similar text generation capabilities, such as "write a birthday wish", do not need to be accessed by the search engine.
[0045] In an exemplary embodiment, the user query text may be directly sent to the search engine, or keywords may be extracted from the user query text and the spoken keywords may be rewritten into written language keywords before being sent to the search engine.
[0046] For example, the user query text "What good movies are coming out soon?" is rewritten into the search keyword "What good movies are coming out recently?"
[0047] In an exemplary embodiment, before sending the user query text to the search engine, a security audit is first performed on the user query text, and the user query text that passes the security audit is sent to the search engine. Some illegal and non-compliant user query texts will not be sent to the search engine.
[0048] In an exemplary embodiment, the search results obtained based on the user query text may come from the same search engine or from multiple search engines. If the search results come from multiple search engines, each search engine may return the same number of search results or a different number of search results, and this application is not limited to this. The search results returned by each search engine may be stored for subsequent unified sorting.
[0049] Step 102 , determining whether the user query text is a time-sensitive query text, if so, executing step 103 , otherwise, executing step 104 .
[0050] In exemplary embodiments, a time-sensitive query text refers to a query text that is time-sensitive, meaning the user is looking for the latest information or information from a specific time period. For example, "latest car purchase subsidy policy" is a time-sensitive query text. However, search engines may search based on keywords and semantics, and may find car purchase subsidy policies from last year or in the past. Therefore, before performing unified sorting, it is necessary to determine the timeliness of the user's query text to provide a timeliness ranking basis for the unified sorting.
[0051] In some embodiments, as Figure 2 As shown, step 102 includes:
[0052] Step 201 : Based on the semantics of the user query text, determine whether the user query text is a time-sensitive query text, and obtain a first determination result.
[0053] In exemplary embodiments, some user queries do not include a specific time. Therefore, whether a user query is a time-sensitive query can be determined purely from a semantic perspective. For example, a user query containing words such as "recently" or "latest" is semantically considered a time-sensitive query. For example, the user query "latest progress in artificial intelligence" contains the word "latest" and is semantically considered a time-sensitive query.
[0054] In some embodiments, step 201 may include: determining whether the user query text contains timeliness semantic keywords; in the case where the user query text contains timeliness semantic keywords, determining that the first judgment result includes that the user query text is a timeliness query text; in the case where the user query text does not contain timeliness semantic keywords, determining that the first judgment result includes that the user query text is not a timeliness query text.
[0055] For example, timeliness semantic keywords may include "recently", "latest", etc., and may also include other words, which are not listed here one by one.
[0056] In other embodiments, Figure 3 As shown, step 201 includes:
[0057] Step 301: Input the user query text into a temporal semantic recognition model to obtain a classification result of the user query text.
[0058] The classification result is used to indicate whether the semantics of the user's query text is relevant to timeliness.
[0059] Among them, the timeliness semantic recognition model is trained based on sample data, and the sample data includes positive sample text, a first label corresponding to the positive sample text, negative sample text, and a second label corresponding to the negative sample text. The first label is used to indicate that the semantics of the positive sample text is related to timeliness, and the second label is used to indicate that the semantics of the negative sample text is not related to timeliness.
[0060] In an exemplary embodiment, the classification results of the user query text may include whether the semantics of the user query text is relevant to timeliness, or whether the semantics of the user query text is not relevant to timeliness. It may also include the probability that the semantics of the user query text is relevant to timeliness. This application does not limit this.
[0061] In an exemplary embodiment, the time-sensitive semantic recognition model can be a large model, and the model training can be performed by fine-tuning the large model, and deployed through a model fine-tuning platform such as a Maas (Model as a Service) platform.
[0062] Step 302: Based on the classification result of the user query text, determine whether the user query text is a time-sensitive query text to obtain a first determination result.
[0063] By labeling the training data with labels indicating whether the semantics are relevant to timeliness and training the model with the labeled training data, the model will be able to determine whether the semantics of the user query text is relevant to timeliness from a semantic perspective. This will enable the model to determine whether the semantics of the user query text is relevant to timeliness from a semantic perspective, thereby improving the accuracy of determining whether the user query text is a time-sensitive query text.
[0064] Step 202: Based on the release time of each search result, determine whether the user query text is a time-sensitive query text to obtain a second determination result.
[0065] In an exemplary embodiment, the publishing time of the search results refers to the time when the document containing the search results can be viewed by the public for the first time.
[0066] In some embodiments, as Figure 4 As shown, step 202 includes:
[0067] Step 401: Determine the first number of first search results in each search result; wherein the first search result includes a publishing time.
[0068] In an exemplary embodiment, the publishing time of each search result is extracted, the search results for which the publishing time cannot be extracted are deleted, and a first number of first search results for which the publishing time can be extracted is counted, and the first number can be recorded as M.
[0069] Step 402: Determine the second number of second search results in each search result; wherein the difference obtained by subtracting the publishing time of the second search result from the current time is less than a preset time length.
[0070] In an exemplary embodiment, the preset duration may be two days or other preset durations, and this application does not limit this. The second number may be denoted as N.
[0071] Step 403: Determine a ratio obtained by dividing the second quantity by the first quantity.
[0072] In an exemplary embodiment, the first number may be denoted as M, the second number may be denoted as N, and a ratio obtained by dividing the second number by the first number may be N / M.
[0073] Step 404 , determining whether the ratio is greater than a preset ratio, if so, executing step 405 , otherwise, executing step 406 .
[0074] In an exemplary embodiment, the preset ratio may be 20%, or may be other preset ratios, and this application is not limited thereto.
[0075] Step 405: Determine whether the second judgment result includes that the user query text is a time-sensitive query text.
[0076] Step 406: Determine that the second judgment result includes that the user query text is not a time-sensitive query text.
[0077] When the ratio of the second number divided by the first number is greater than the preset ratio, it can be considered that the user query text has generated recent search results and is a time-sensitive query text, which can improve the accuracy of determining whether the user query text is a time-sensitive query text.
[0078] In other embodiments, step 202 may include: determining a second number of second search results in each search result; wherein, the difference obtained by subtracting the release time of the second search result from the current time is less than a preset time length; when the second number is greater than the preset number, determining that the second judgment result includes that the user query text is a time-sensitive query text; when the second number is less than or equal to the preset number, determining that the second judgment result includes that the user query text is not a time-sensitive query text.
[0079] It is also possible to determine whether the user query text has recent search results by whether the second number is greater than the preset number, and then determine whether the user query text is a time-sensitive query text, which can improve the accuracy of determining whether the user query text is a time-sensitive query text.
[0080] Step 203: Based on the user query text and the hot event, determine whether the user query text is a time-sensitive query text to obtain a third determination result.
[0081] In an exemplary embodiment, a hot event may refer to a recent hot event obtained by crawling or manual configuration, for example, a hot event may be an event in a hot search, a hot event that is being discussed recently, and new information may be generated recently.
[0082] In some embodiments, as Figure 5 As shown, step 203 includes:
[0083] Step 501: extract entity words from hot events to obtain various hot entity words.
[0084] Hot events are being discussed recently, and new information will be generated recently. Hot entity words extracted from hot events will also have new information generated recently.
[0085] Step 502: extract entity words from the user query text to obtain each query entity word.
[0086] Step 503: Determine the semantic similarity between each query entity word and each hot entity word.
[0087] Step 504 , determining whether there is at least one query entity word in each query entity word and at least one hot entity word whose semantic similarity is greater than a preset similarity; if so, executing step 505 ; otherwise, executing step 506 .
[0088] Step 505: Determine whether the third judgment result includes that the user query text is a time-sensitive query text.
[0089] Step 506: Determine whether the third judgment result includes that the user query text is not a time-sensitive query text.
[0090] When the semantic similarity between at least one query entity word and at least one hot entity word in each query entity word is greater than the preset similarity, it is considered that at least one query entity word and the hot entity word in the user query text have a successful semantic match, and it is determined that the user is asking for information related to the hot entity, which is a time-sensitive intention, and can improve the accuracy of judging whether the user query text is a time-sensitive query text.
[0091] In other embodiments, step 203 may include: determining the semantic similarity between the user query text and the hot event; when the semantic similarity between the user query text and the hot event is greater than the similarity threshold, determining that the third judgment result includes that the user query text is a timeliness query text; when the semantic similarity between the user query text and the hot event is less than or equal to the similarity threshold, determining that the third judgment result includes that the user query text is not a timeliness query text.
[0092] It is also possible to directly determine the semantic similarity between the user query text and the hot events without extracting entity words, and match the user query text and the hot events as a whole. If the match is successful, it is determined that the user is asking for information related to the hot event, which is a timely intention, and can improve the accuracy of judging whether the user query text is a timely query text.
[0093] Step 204 , determine whether there is at least one judgment result indicating that the user query text is a time-sensitive query text among the first judgment result, the second judgment result, and the third judgment result. If so, execute step 205 ; otherwise, execute step 206 .
[0094] Step 205: Determine whether the user query text is a time-sensitive query text.
[0095] Step 206: Determine whether the user query text is a time-sensitive query text.
[0096] The timeliness of the user query text is judged from three different dimensions: the semantics of the user query text, the release time of each search result, and the matching between the user query text and hot events. Among the first judgment result, the second judgment result and the third judgment result, if there is at least one judgment result indicating that the user query text is a timely query text, that is, when any dimension is hit, the user query text is considered to be a timely query text. Compared with the solution of judging the timeliness of the user query text from a single dimension, it can reduce the occurrence of user query texts belonging to timely query texts being misjudged as non-timely query texts, and can improve the accuracy of judging whether the user query text is a timely query text.
[0097] In other embodiments, step 102 may include determining whether the user query text is a time-sensitive query text based on the semantics of the user query text; step 102 may also include determining whether the user query text is a time-sensitive query text based on the time when each search result is published; step 102 may also include determining whether the user query text is a time-sensitive query text based on the user query text and hot events. The timeliness of the user query text can be directly determined from a single dimension, and this application is not limited to this.
[0098] Step 103: sort the search results based on the relevance score and timeliness score of the search results to obtain a sorting result.
[0099] Among them, the sorting results are used to assist the large model in generating reply text corresponding to the user query text.
[0100] In an exemplary embodiment, the large model can generate a reply text corresponding to the user query text based on the search results that are ranked at the top of the sorted results.
[0101] In an exemplary embodiment, steps 101 to 104 in the present application are implemented outside the large model and are independent of the large model.
[0102] In an exemplary embodiment, the relevance score is used to indicate the relevance between the search results and the user query text. A higher relevance score indicates a stronger relevance between the search results and the user query text.
[0103] In an exemplary embodiment, the relevance score between the search results and the user query text can be obtained by using a cross encoder or a dual encoder. The relevance score obtained by using a cross encoder is more accurate.
[0104] In an exemplary embodiment, a higher timeliness score indicates that the publishing time of the search result is closer to the current time.
[0105] In an exemplary embodiment, the timeliness score may be a score calculated based on a difference between the current time and the publishing time.
[0106] When the user query text is a time-sensitive query text, it indicates that the timeliness of the search results can affect the accuracy of the search results. In the process of sorting the search results, the relevance score and timeliness score of the search results are comprehensively considered. On the premise of ensuring the relevance sorting, the search results with strong timeliness are further given priority, which can improve the accuracy of the sorting results. The sorting results are used to assist the large model in generating the reply text corresponding to the user query text, further improving the accuracy of the reply text generated by the large model corresponding to the user query text.
[0107] In some embodiments, as Figure 6 As shown, step 103 includes:
[0108] Step 601 : Based on the relevance score of each search result, the search results are divided according to each preset relevance score interval to obtain each search result subset.
[0109] A subset of search results corresponds to a preset relevance score range.
[0110] Step 602 : sorting the search result subsets corresponding to the respective preset relevance score intervals according to the arrangement order of the respective preset relevance score intervals to obtain a subset sorting result.
[0111] Step 603 : sorting the search results in the search result subset based on the timeliness score of each search result in the search result subset to obtain a sorting result within the subset.
[0112] In an exemplary embodiment, step 603 may include: within the search result subset, sorting the search results in the search result subset based solely on the timeliness score of each search result in the search result subset, thereby obtaining an internal sorting result within the subset; step 603 may also include: within the search result subset, sorting the search results in the search result subset based on both the timeliness score and the authority score of each search result in the search result subset, thereby obtaining an internal sorting result within the subset. Other implementations are also possible, and this application is not limited thereto.
[0113] Step 604: Obtain a sorting result based on the subset sorting result and the sub-set internal sorting result.
[0114] For example, the preset relevance score intervals include [1, 0.8), [0.8, 0.6), [0.6, 0.4), [0.4, 0.2), and [0.2, 0]. The relevance scores of the search results are: the relevance score of search result A is 0.9, the relevance score of search result B is 0.85, the relevance score of search result C is 0.7, the relevance score of search result D is 0.8, and the relevance score of search result E is 0.5. The preset relevance score intervals are sorted from high to low according to the relevance score, as follows: [1, 0.8), [0.8, 0.6), [0.6, 0.4), [0.4, 0.2), and [0.2, 0]. Based on the relevance score of each search result, the search results are divided according to each preset relevance score interval to obtain the search result subset {search result A, search result B} corresponding to the preset relevance score interval [1,0.8), the search result subset {search result C, search result D} corresponding to the preset relevance score interval [0.8,0.6), and the search result subset {search result E} corresponding to the preset relevance score interval [0.6,0.4). According to the arrangement order of each preset relevance score interval, the subset sorting results are {search result A, search result B}, {search result C, search result D}, and {search result E}. The timeliness score of search result A is less than the timeliness score of search result B, and the timeliness score of search result C is greater than the timeliness score of search result D. The internal sorting results of the subsets are {search result B, search result A} and {search result C, search result D}, and the final sorting result is {search result B, search result A, search result C, search result D, search result E}.
[0115] First, sort the search result subsets corresponding to each preset relevance score interval according to the arrangement order of each preset relevance score interval, and then sort the search results in the search result subset based on the timeliness score of each search result in the search result subset, to ensure that the relevance sorting is not affected by timeliness, and put the latest but irrelevant content in the front. On the premise of ensuring the relevance sorting, further give priority to displaying the timeliness content, which can improve the accuracy of the sorting results and further improve the accuracy of the reply text corresponding to the user query text generated by the large model.
[0116] In other embodiments, step 103 may include: performing a weighted summation based on the relevance score and timeliness score of each search result to obtain a comprehensive score for each search result; and sorting each search result based on the comprehensive score of each search result to obtain a sorted result. Sorting each search result based on the relevance score and timeliness score of each search result to obtain a sorted result may also be achieved by weighted summation, and this application is not limited thereto.
[0117] Step 104: sort the search results based on their relevance scores to obtain sorting results.
[0118] When the user query text is not a time-sensitive query text, it indicates that the timeliness of the search results will not affect the accuracy of the search results. In the process of sorting the search results, the relevance score of the search results is considered, and the timeliness score of the search results is not considered. Reducing the adverse impact of the timeliness score on the sorting results can improve the accuracy of the sorting results and further improve the accuracy of the reply text generated by the large model corresponding to the user query text.
[0119] In some embodiments, as Figure 7 As shown, step 104 includes:
[0120] Step 701: Obtain the authority level of the data source of each search result.
[0121] For example, the authority levels may include high authority, medium authority, and low authority, and may also include other level division methods, which are not limited in this application.
[0122] In some embodiments, as Figure 8 As shown, step 701 includes:
[0123] Step 801 : According to the fields to which the historical user query texts belong, statistics are collected on the data sources of the historical search results corresponding to the historical user query texts to obtain the corresponding quantities of the data sources in the fields.
[0124] In an exemplary embodiment, the data source may refer to first-level domain name information, that is, a first-level site.
[0125] Step 802 : sorting the data sources in each field in descending order based on the quantity corresponding to each data source in each field, and obtaining the sorting results of the data sources corresponding to each field.
[0126] Step 803: Based on the ranking results of the data sources corresponding to each field, determine the authority level corresponding to each data source in each field.
[0127] For example, in the data source sorting results in domain 1, Top1~TopN are high-authority sites with a high authority level; TopN~Top2N are medium-authority sites with a medium authority level; Top2N~others are low-authority sites with a low authority level.
[0128] Step 804 : determining the authority level corresponding to each data source in the target domain based on the target domain to which the user query text belongs and the authority level corresponding to each data source in each domain.
[0129] Step 805 : Determine the authority level of the data source of each search result based on the authority level corresponding to each data source in the target domain.
[0130] Traditional authoritative sites are manually sorted, which is time-consuming and labor-intensive. This application can dynamically adjust the authority level of data sources based on historical search results without manual intervention. It can also dynamically adjust the authority level of data sources regularly to ensure the accuracy of the authority level of data sources.
[0131] Step 702: Determine the authority score of each search result based on the authority level of the data source of each search result.
[0132] For example, the authority score of a search result whose data source has a high authority level is 0.1 points, the authority score of a search result whose data source has a medium authority level is 0.05 points, and the authority score of a search result whose data source has a low authority level is 0 points. If a search result has multiple data sources, and the authority levels of the multiple data sources include multiple authority levels, the authority score of the higher authority level is determined as the authority score of the search result. For example, a search result has two data sources, one data source has a high authority level, and the other data source has a medium authority level. The authority score of this search result is 0.1 points. The authority score can also be calculated in other ways, and this application does not limit this.
[0133] Step 703: sort the search results based on their relevance scores and authority scores to obtain sorting results.
[0134] In an exemplary embodiment, step 703 may include: performing a weighted summation based on the relevance score and authority score of each search result to obtain a comprehensive score for each search result; and ranking each search result based on the comprehensive score to obtain a ranking result. The relevance score is weighted more heavily than the authority score, with relevance being the most important. The ranking is based on the relevance score, and the final score is adjusted based on the authority score, based on the relevance.
[0135] Taking relevance and authority scores into consideration during ranking can improve the accuracy of the ranking results and further improve the accuracy of the response text generated by the large model corresponding to the user query text.
[0136] In other embodiments, step 104 may also include: sorting the search results based only on the relevance scores of the search results to obtain sorting results.
[0137] In some embodiments, as Figure 9 As shown, before sorting the search results in step 103 and obtaining the sorted results, or before sorting the search results in step 104 and obtaining the sorted results, the search sorting method further includes:
[0138] Step 901: Obtain the authority level of the data source of each search result.
[0139] In an exemplary embodiment, step 901 may implement dynamic adjustment of the authority level of the data source according to steps 801 to 805 .
[0140] Step 902: for any search result, determine the number of data sources corresponding to each search result at each authority level.
[0141] Step 903 : A trust score of any search result is obtained by performing weighted summation based on the number of data sources corresponding to each authority level and the trust weight corresponding to each authority level.
[0142] For example, if search result A has m high-authority sites, n medium-authority sites, and k low-authority sites, then the trust score of search result A is m*a+n*b+k*c, where a, b, and c are the weight values corresponding to high-, medium-, and low-authority sites, respectively. They can be set according to actual conditions, where a is greater than b and b is greater than c.
[0143] Step 904: Determine untrustworthy search results from the search results based on the trust scores of the search results.
[0144] In an exemplary embodiment, when the trust score of search result A exceeds the trust score of search result B by 50%, search result A is considered a trustworthy search result and search result B is considered an untrustworthy search result. Otherwise, both search results A and B are considered trustworthy search results.
[0145] Step 905: Delete untrustworthy search results from each search result.
[0146] In an exemplary embodiment, when searching, it is often found that individual websites have outdated or incorrect information. For example, for "Administrative district planning of City A", there will be multiple different search results because the administrative districts of City A have been adjusted. Therefore, it is necessary to perform consistency verification on the search results to remove outdated and incorrect search results.
[0147] Before sorting, based on the trust score of each search result, untrustworthy search results are identified from each search result, and untrustworthy search results are deleted from each search result. Incorrect and outdated information on the Internet is filtered to ensure the accuracy of the search results. This can improve the accuracy of the sorting results and further improve the accuracy of the reply text generated by the large model corresponding to the user query text.
[0148] In summary, in this application, each search result obtained by searching based on the user query text is obtained, and it is first determined whether the user query text is a time-sensitive query text. If the user query text is a time-sensitive query text, it indicates that the timeliness of the search results can affect the accuracy of the search results. In the process of sorting the search results, the relevance score and timeliness score of the search results are comprehensively considered. Under the premise of ensuring the relevance sorting, the search results with strong timeliness are further given priority, which can improve the accuracy of the sorting results. The sorting results are used to assist the large model in generating the reply text corresponding to the user query text, further improving the accuracy of the reply text generated by the large model for the user query text. In the case that the user query text is not a time-sensitive query text, it indicates that the timeliness of the search results does not affect the accuracy of the search results. In the process of sorting the search results, the relevance score of the search results is considered, and the timeliness score of the search results is not considered. This reduces the adverse effect of the timeliness score on the sorting results, can improve the accuracy of the sorting results, and further improves the accuracy of the reply text generated by the large model for the user query text.
[0149] Exemplary devices
[0150] Accordingly, the embodiment of the present application also provides a search and sorting device, such as Figure 10 As shown, the search and sorting device includes:
[0151] Search unit 1001, used to obtain search results based on the user query text;
[0152] A judging unit 1002 is configured to judge whether the user query text is a time-sensitive query text;
[0153] A first ranking unit 1003 is configured to, when the user query text is a timeliness query text, rank the search results based on the relevance score and timeliness score of each search result to obtain a ranking result;
[0154] A second ranking unit 1004 is configured to, when the user query text is not a time-sensitive query text, rank the search results based on the relevance scores of the search results to obtain a ranking result;
[0155] The ranking result is used to assist the large model in generating a reply text corresponding to the user query text.
[0156] Optionally, the judging unit 1002 includes:
[0157] A first judgment subunit is configured to judge whether the user query text is a time-sensitive query text based on the semantics of the user query text, and obtain a first judgment result;
[0158] A second judgment subunit is configured to judge whether the user query text is a time-sensitive query text based on the release time of each search result, and obtain a second judgment result;
[0159] A third judgment subunit is configured to judge whether the user query text is a time-sensitive query text based on the user query text and the hot event, and obtain a third judgment result;
[0160] a first processing subunit, configured to determine that the user query text is a time-sensitive query text if at least one judgment result among the first judgment result, the second judgment result, and the third judgment result indicates that the user query text is a time-sensitive query text;
[0161] The second processing subunit is configured to determine that the user query text is not a time-sensitive query text when the first judgment result, the second judgment result, and the third judgment result all indicate that the user query text is not a time-sensitive query text.
[0162] Optionally, the first judgment subunit is specifically configured to:
[0163] Inputting the user query text into a timeliness semantic recognition model to obtain a classification result of the user query text; wherein the classification result is used to indicate whether the semantics of the user query text is related to timeliness;
[0164] Based on the classification result of the user query text, determining whether the user query text is a time-sensitive query text to obtain a first determination result;
[0165] In which, the timeliness semantic recognition model is trained based on sample data, and the sample data includes positive sample text, a first label corresponding to the positive sample text, negative sample text, and a second label corresponding to the negative sample text. The first label is used to indicate that the semantics of the positive sample text is related to timeliness, and the second label is used to indicate that the semantics of the negative sample text is not related to timeliness.
[0166] Optionally, the second judgment subunit is specifically configured to:
[0167] Determining a first number of first search results in each search result; wherein the first search result includes a publishing time;
[0168] Determining a second number of second search results in each search result; wherein a difference obtained by subtracting a publishing time of the second search result from a current time is less than a preset time length;
[0169] determining a ratio of the second amount divided by the first amount;
[0170] In a case where the ratio is greater than a preset ratio, determining that the second judgment result includes that the user query text is a time-sensitive query text;
[0171] When the ratio is less than or equal to the preset ratio, it is determined that the second judgment result includes that the user query text is not a time-sensitive query text.
[0172] Optionally, the third judgment subunit is specifically configured to:
[0173] Extract entity words from hot events to obtain various hot entity words;
[0174] Extract entity words from the user query text to obtain each query entity word;
[0175] Determine the semantic similarity between each query entity word and each hot entity word;
[0176] In the case where the semantic similarity between at least one query entity word and at least one hot entity word among the query entity words is greater than a preset similarity, determining that the third judgment result includes that the user query text is a timeliness query text;
[0177] When the semantic similarity between any query entity word and any hot entity word in the query entity words is less than or equal to the preset similarity, it is determined that the third judgment result includes that the user query text is not a time-sensitive query text.
[0178] Optionally, the first sorting unit 1003 is specifically configured to:
[0179] Based on the relevance score of each search result, the search results are divided according to each preset relevance score interval to obtain each search result subset; wherein one search result subset corresponds to one preset relevance score interval;
[0180] Sorting the search result subsets corresponding to the respective preset relevance score intervals according to the arrangement order of the respective preset relevance score intervals to obtain a subset sorting result;
[0181] Within the search result subset, sorting the search results in the search result subset based on the timeliness score of each search result in the search result subset to obtain a sorting result within the subset;
[0182] A sorting result is obtained based on the subset sorting result and the internal sorting result of the subset.
[0183] Optionally, the second sorting unit 1004 is specifically configured to:
[0184] Get the authority level of the data source for each search result;
[0185] Determining an authority score for each search result based on the authority level of the data source of each search result;
[0186] Based on the relevance score and authority score of each search result, each search result is sorted to obtain a sorted result.
[0187] Optionally, the search ranking device further includes:
[0188] An acquisition unit, used to obtain the authority level of the data source of each search result;
[0189] A first processing unit is configured to determine, for any search result, the number of data sources corresponding to the search result at each authority level;
[0190] A second processing unit is configured to perform weighted summation based on the number of data sources corresponding to each authority level of the arbitrary search result and the trust weight corresponding to each authority level to obtain a trust score for the arbitrary search result;
[0191] a third processing unit, configured to determine untrustworthy search results from the search results based on the trust scores of the search results;
[0192] The deleting unit is configured to delete the untrustworthy search results from each search result.
[0193] Optionally, the acquisition unit is specifically configured to:
[0194] According to the fields to which each historical user query text belongs, statistics are collected on the data sources of the historical search results corresponding to each historical user query text, and the corresponding quantity of each data source in each field is obtained;
[0195] Based on the corresponding quantity of each data source in each field, the data sources in each field are sorted in descending order to obtain the corresponding data source sorting results for each field;
[0196] Based on the ranking results of the data sources corresponding to each field, determine the authority level of each data source in each field;
[0197] Determining the authority level corresponding to each data source in the target domain based on the target domain to which the user query text belongs and the authority level corresponding to each data source in each domain;
[0198] Based on the authority level corresponding to each data source in the target domain, the authority level of the data source of each search result is determined.
[0199] The search ranking device provided in this embodiment is based on the same concept as the search ranking method provided in the above embodiments of this application. It can execute the search ranking method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of executing the search ranking method. For technical details not fully described in this embodiment, please refer to the specific processing content of the search ranking method provided in the above embodiments of this application and will not be repeated here.
[0200] The functions implemented by the above-mentioned search unit 1001, judgment unit 1002, first sorting unit 1003 and second sorting unit 1004 can be respectively implemented by the same or different processors, which is not limited in the embodiment of the present application.
[0201] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, and the memory can be a memory within the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits. The functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units can be realized by designing the logical relationships between the components within the circuit. For another example, in another implementation, the hardware circuit can be implemented by a PLD. For example, an FPGA can include a large number of logic gate circuits. The connection relationships between the logic gate circuits are configured through a configuration file to realize the functions of some or all of the above units. All units of the above devices can be implemented entirely in the form of a processor calling software, or entirely in the form of hardware circuits, or partially in the form of a processor calling software, with the remaining parts implemented in the form of hardware circuits.
[0202] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP. In another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit may be fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0203] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0204] In addition, the various units in the above apparatus may be fully or partially integrated together, or may be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the apparatus. The at least one processor may be of different types, such as a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0205] Exemplary electronic devices
[0206] An embodiment of the present application provides an electronic device, see Figure 11 As shown, the device includes:
[0207] Memory 200 and processor 210;
[0208] The memory 200 is connected to the processor 210 and is used to store programs;
[0209] The processor 210 is configured to implement the search and sorting method disclosed in any one of the above embodiments by running the program stored in the memory 200 .
[0210] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0211] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are interconnected via a bus.
[0212] A bus may include a pathway that transfers information between components of a computer system.
[0213] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like. It can also be an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware components.
[0214] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0215] Memory 200 stores programs that implement the technical solutions of the present invention and may also store an operating system and other key services. Specifically, the programs may include program code, which includes computer operating instructions. More specifically, memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, and the like.
[0216] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0217] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speakers, etc.
[0218] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0219] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any search and ranking method provided in the above embodiments of the present application.
[0220] Exemplary Question Answering System
[0221] The present application also provides a question-answering system. Figure 12 As shown, the question-answering system includes:
[0222] Large model 1201, search engine 1202 and electronic device 1203;
[0223] Large model 1201 is configured to receive a user query text, and if it is determined that a reply to the user query text is required via an information query, send the user query text to search engine 1202, send each search result returned by search engine 1202 to electronic device 1203, and generate a reply text corresponding to the user query text based on the feedback from electronic device 1203;
[0224] The search engine 1202 is configured to, upon receiving a user query text sent by the large model 1201 , perform a search based on the user query text, obtain search results, and feed the search results back to the large model 1201 ;
[0225] The electronic device 1203 is used to process the search results sent by the large model 1201 according to the search ranking method disclosed in any of the above embodiments when receiving the search results, and feed back the processing results to the large model 1201.
[0226] For technical details not fully described in this embodiment, please refer to the specific processing content of the search ranking method provided in the above embodiment of this application and the specific content of the electronic device provided in the above embodiment of this application, and will not be repeated here.
[0227] Exemplary computer program products and storage media
[0228] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the search ranking method according to various embodiments of the present application described in any of the above-mentioned embodiments of this specification.
[0229] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0230] In addition, an embodiment of the present application may also be a storage medium having a computer program stored thereon, and the computer program is used by a processor to execute the steps of the search and ranking method according to various embodiments of the present application described in any of the above embodiments of this specification, specifically the following steps:
[0231] Step 101: Obtain search results based on a user query text.
[0232] Step 102 , determining whether the user query text is a time-sensitive query text, if so, executing step 103 , otherwise, executing step 104 .
[0233] Step 103: sort the search results based on the relevance score and timeliness score of the search results to obtain a sorting result.
[0234] Step 104: sort the search results based on their relevance scores to obtain sorting results.
[0235] Among them, the sorting results are used to assist the large model in generating reply text corresponding to the user query text.
[0236] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0237] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0238] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0239] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be merged, divided, and deleted according to actual needs.
[0240] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0241] The modules or submodules described as separate components may or may not be physically separate, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the purpose of this embodiment according to actual needs.
[0242] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.
[0243] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0244] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software executed by a processor, or a combination of the two. The software may be stored in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0245] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0246] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A search ranking method, characterized in that: include: Obtaining search results based on the user's query text; Determining whether the user query text is a time-sensitive query text; In the case where the user query text is a time-sensitive query text, based on the relevance scores of the respective search results, the search results are divided according to the respective preset relevance score intervals to obtain respective search result subsets; wherein, one search result subset corresponds to one preset relevance score interval; according to the arrangement order of the respective preset relevance score intervals, the search result subsets corresponding to the respective preset relevance score intervals are sorted to obtain a subset sorting result; within the search result subsets, based on the timeliness scores of the respective search results in the search result subsets, the search results in the search result subsets are sorted to obtain a subset internal sorting result; and a sorting result is obtained based on the subset sorting result and the subset internal sorting result; In the case where the user query text is not a time-sensitive query text, sorting the search results based on the relevance scores of the search results to obtain a sorting result; The ranking result is used to assist the large model in generating a reply text corresponding to the user query text.
2. The search ranking method according to claim 1, characterized in that: The determining whether the user query text is a time-sensitive query text includes: Based on the semantics of the user query text, determining whether the user query text is a time-sensitive query text, and obtaining a first determination result; Based on the release time of each search result, determining whether the user query text is a time-sensitive query text to obtain a second determination result; Based on the user query text and the hot event, determining whether the user query text is a time-sensitive query text to obtain a third determination result; If at least one of the first judgment result, the second judgment result, and the third judgment result indicates that the user query text is a time-sensitive query text, determining that the user query text is a time-sensitive query text; In a case where the first judgment result, the second judgment result, and the third judgment result all indicate that the user query text is not a time-sensitive query text, it is determined that the user query text is not a time-sensitive query text.
3. The search ranking method according to claim 2, characterized in that: The determining, based on the semantics of the user query text, whether the user query text is a time-sensitive query text to obtain a first determination result includes: Inputting the user query text into a timeliness semantic recognition model to obtain a classification result of the user query text; wherein the classification result is used to indicate whether the semantics of the user query text is related to timeliness; Based on the classification result of the user query text, determining whether the user query text is a time-sensitive query text to obtain a first determination result; In which, the timeliness semantic recognition model is trained based on sample data, and the sample data includes positive sample text, a first label corresponding to the positive sample text, negative sample text, and a second label corresponding to the negative sample text. The first label is used to indicate that the semantics of the positive sample text is related to timeliness, and the second label is used to indicate that the semantics of the negative sample text is not related to timeliness.
4. The search ranking method according to claim 2, wherein: The determining, based on the release time of each search result, whether the user query text is a time-sensitive query text to obtain a second determination result includes: Determining a first number of first search results in each search result; wherein the first search result includes a publishing time; Determining a second number of second search results in each search result; wherein a difference obtained by subtracting a publishing time of the second search result from a current time is less than a preset time length; determining a ratio of the second amount divided by the first amount; In a case where the ratio is greater than a preset ratio, determining that the second judgment result includes that the user query text is a time-sensitive query text; When the ratio is less than or equal to the preset ratio, it is determined that the second judgment result includes that the user query text is not a time-sensitive query text.
5. The search ranking method according to claim 2, characterized in that: The determining, based on the user query text and the hot event, whether the user query text is a time-sensitive query text to obtain a third determination result includes: Extract entity words from hot events to obtain various hot entity words; Extract entity words from the user query text to obtain each query entity word; Determine the semantic similarity between each query entity word and each hot entity word; In the case where the semantic similarity between at least one query entity word and at least one hot entity word among the query entity words is greater than a preset similarity, determining that the third judgment result includes that the user query text is a timeliness query text; When the semantic similarity between any query entity word and any hot entity word in the query entity words is less than or equal to the preset similarity, it is determined that the third judgment result includes that the user query text is not a time-sensitive query text.
6. The search ranking method according to claim 1, characterized in that: The step of sorting the search results based on the relevance scores of the search results to obtain sorting results includes: Get the authority level of the data source for each search result; Determining an authority score for each search result based on the authority level of the data source of each search result; Based on the relevance score and authority score of each search result, each search result is sorted to obtain a sorted result.
7. The search ranking method according to claim 1, characterized in that: Before sorting the search results and obtaining the sorted results, the method further includes: Get the authority level of the data source for each search result; For any search result, determine the number of data sources corresponding to the search result at each authority level; A trust score for any search result is obtained by performing a weighted sum based on the number of data sources corresponding to each authority level and the trust weight corresponding to each authority level; determining untrustworthy search results from the search results based on the trust scores of the search results; The untrustworthy search results are deleted from each search result.
8. The search ranking method according to claim 6 or 7, characterized in that: The authority level of the data source of each search result is obtained, including: According to the fields to which each historical user query text belongs, statistics are collected on the data sources of the historical search results corresponding to each historical user query text, and the corresponding quantity of each data source in each field is obtained; Based on the corresponding quantity of each data source in each field, the data sources in each field are sorted in descending order to obtain the corresponding data source sorting results for each field; Based on the ranking results of the data sources corresponding to each field, determine the authority level of each data source in each field; Determining the authority level corresponding to each data source in the target domain based on the target domain to which the user query text belongs and the authority level corresponding to each data source in each domain; Based on the authority level corresponding to each data source in the target domain, the authority level of the data source of each search result is determined.
9. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the search and sorting method according to any one of claims 1 to 8 by running the program in the memory.
10. A question-answering system, characterized in that: include: large models, search engines, and electronic devices; The large model is configured to receive a user query text, and when determining that the user query text needs to be replied to through an information query, send the user query text to the search engine, send each search result fed back by the search engine to the electronic device, and generate a reply text corresponding to the user query text based on the feedback from the electronic device; The search engine is configured to, upon receiving a user query text sent by the large model, perform a search based on the user query text, obtain search results, and feed the search results back to the large model; The electronic device is configured to process the search results sent by the large model according to the search ranking method according to any one of claims 1 to 8 upon receipt thereof, and feed back the processing results to the large model.
Citation Information
Patent Citations
Information search methods and devices and timeliness query term recognition methods and devices
CN107180093B
Text processing method and device, electronic equipment and computer readable storage medium
CN116340467A