Sorting method and device, medium, electronic equipment and program product
By introducing exact match, substring match, and scattered match types into the search system and calculating the target match degree, the problems of inaccurate ranking results and high cost in the existing technology are solved, and more accurate ranking results and lower search costs are achieved.
Patent Information
- Application Number
- CN202511822167.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-01-02
AI Technical Summary
Existing search systems lack strict matching type order relationships during data search and sorting, resulting in inaccurate sorting results and high search costs. There is also a lack of unified matching features between multiple sorting layers, leading to inconsistencies between layers and sorting instability.
By obtaining the matching type between the text field and the question to be searched, including exact match, substring match and scattered match, the target matching degree is calculated, and the candidate set is sorted based on this. A strict matching type order relationship is introduced to ensure the accuracy of the sorting results.
It achieves more accurate sorting results, reduces search costs, solves the problems of inaccurate sorting results and high costs in existing technologies, and improves the overall efficiency of the search system.
Smart Images

Figure CN121255871A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular, to a ranking method, device, medium, electronic equipment and program product. BACKGROUND
[0002] With the development of big data, various types of data are growing explosively, and the information data accessible to people is growing continuously, i.e., a search system can obtain massive data when receiving a search question input by a user. Therefore, how to search and rank data based on the search question input by the user to provide more accurate ranking results for the user is a technical problem to be solved urgently. SUMMARY
[0003] This summary is provided to introduce a selection of concepts, which will be described with greater specificity in the detailed description section. This summary does not intend to identify key or essential features of the claimed technology, nor is it intended for use in determining the scope of the claimed technology.
[0004] In a first aspect, the present disclosure provides a ranking method, the method comprising: receiving a search question, obtaining a candidate set corresponding to the search question, the candidate set comprising at least one text field; determining a matching type between the text field and the search question, the matching type comprising a complete matching type, a substring matching type and a scattered matching type; obtaining a target matching degree between the text field and the search question based on the matching type; ranking the at least one text field according to the target matching degree to obtain a ranking result.
[0005] In a second aspect, the present disclosure provides a ranking device, the device comprising: a first obtaining module configured to receive a search question, and obtain a candidate set corresponding to the search question, the candidate set comprising at least one text field; a determining module configured to determine a matching type between the text field and the search question, the matching type comprising a complete matching type, a substring matching type and a scattered matching type; a second obtaining module configured to obtain a target matching degree between the text field and the search question based on the matching type; a ranking module configured to rank the at least one text field according to the target matching degree to obtain a ranking result.
[0006] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program which, when executed by a processing apparatus, implements the steps of the method of the first aspect.
[0007] In a fourth aspect, the present disclosure provides an electronic device comprising: a storage device having stored thereon a computer program; a processing apparatus configured to execute the computer program stored in the storage device to implement the steps of the method of the first aspect.
[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0009] Through the above technical solution, in the case of receiving a to-be-searched question, a candidate set corresponding to the to-be-searched question is obtained, the candidate set can include multiple text fields, on this basis, a matching type between the text field and the to-be-searched question is determined, and a target matching degree between the text field and the to-be-searched question is obtained based on the matching type, wherein the matching type includes a complete matching type, a substring matching type and an assignment matching type, and finally the multiple text fields in the candidate set are sorted based on the target matching degree to obtain a sorting result. Since the target matching degree is introduced in the sorting of the candidate set, and the target matching degree is determined through the matching type between the text field and the to-be-searched question, the accuracy of the priority field sorting can be greatly ensured.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which: Figure 1 FIG. 1 is a structural example diagram of a search system according to an example embodiment of the present disclosure.
[0012] Figure 2 FIG. 3 is a flowchart of a sorting method according to an example embodiment of the present disclosure.
[0013] Figure 3 FIG. 4 is an example diagram of a sorting result in a sorting method according to an example embodiment of the present disclosure.
[0014] Figure 4 FIG. 5 is a block diagram of a sorting device according to an example embodiment of the present disclosure.
[0015] Figure 5 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art.
[0017] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0018] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising but not limited to". The term "based on" is "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions are given below.
[0019] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0021] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the permission of the user should be obtained in accordance with relevant laws and regulations.
[0023] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.
[0024] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select “agree” or “disagree” to provide personal information to the electronic device.
[0025] It can be understood that the above notification and obtaining user permission process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0026] At the same time, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0027] In the related art, the search system can use the architecture of a recall layer 101, a sorting layer 102 and a rearrangement layer 103 as shown in Figure 1 to perform data search, wherein the sorting layer 102 can include multiple sorting layers, such as a first sorting layer and a second sorting layer. The first sorting layer can be a coarse sorting layer, and the second sorting layer can be a fine sorting layer.
[0028] For example, the recall layer 101 can perform a recall operation through an inverted vector / cross-modal multi-path after receiving a search question to obtain a large candidate set that can have noise, such as inputting the search question to the recall layer 101, and the output data of the recall layer 101 can be a recall result of the order of ten thousand / one hundred thousand. The sorting layer 102 can perform a light scoring on the candidate set to obtain a smaller candidate list arranged in descending order of “rough score”, such as inputting the candidate list to the sorting layer 102, and the output data of the sorting layer 102 can be a sorting result of the order of one hundred. The rearrangement layer 103 can perform a fine sorting operation through a high-cost model to obtain a search result list with optimal order, such as inputting the sorting result to the rearrangement layer 103, and the output data of the rearrangement layer 103 can be a Top-N result list.
[0029] The related art lacks strict order relations of Exact greater than Substring, Substring greater than BoW and in-class continuous scale in the process of obtaining the sorting result, so that the sorting result is inaccurate. In addition, the related art mainly relies on index-side field division or hard-layer modification in the search process, resulting in high search cost. In addition, there is a lack of unified matching features between multiple sorting layers of the search system, resulting in inconsistency and sorting jitter problems between layers.
[0030] To solve the above problems, the present disclosure proposes a sorting method, device, medium, electronic equipment and program product. In the process of obtaining the sorting result, the matching type between the text field and the search problem is obtained. The target matching degree obtained by using the matching type can more accurately realize sorting. Because the strict order relations of Exact, Substring and BoW are introduced, the sorting result obtained by the embodiment of the present disclosure is more accurate.
[0031] The present disclosure will be further explained and described below in conjunction with the drawings.
[0032] Figure 2 is a flowchart of a sorting method according to an embodiment of the present disclosure. The sorting method can be applied to an electronic device with processing capability, such as a terminal or a server. And the sorting method can be executed by a sorting device, wherein the sorting device can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. Referring to Figure 2 The sorting method can include the following steps.
[0033] In step S210, a search problem is received, and a candidate set corresponding to the search problem is obtained.
[0034] The embodiment of the present disclosure can be applied to a search system, which is based on Figure 1 It can be known that the search system can include a recall layer, multiple sorting layers and a rearrangement layer, etc., that is, the search system can include multiple sorting layers. For example, the multiple sorting layers can include a first sorting layer, a second sorting layer, etc., wherein the first sorting layer can be connected with the recall layer, that is, the input data of the first sorting layer can be the output data of the recall layer.
[0035] As an optional way, in the case of receiving a search problem input by a user, the embodiment of the present disclosure can obtain a candidate set corresponding to the search problem through the recall layer. For example, the candidate set can be a candidate set possibly related to the search problem found from massive data by using an inverted index, vector retrieval, etc. after the recall layer receives the search problem.
[0036] In the embodiments of the present disclosure, the candidate set can include at least one text field and a plurality of other fields, wherein the set composed of at least one text field can be referred to as a priority text attribute set, and the priority text attribute set can include a plurality of text fields, which can also be referred to as priority text fields. For example, the priority text attribute set can include image_title (image title), product_title (product title), and content_title (content title), etc.
[0037] For example, for each data set, the embodiments of the present disclosure can select one field to map to priority_text, that is, the priority_text can include a plurality of text fields, and each text field in the priority_text can be traversed subsequently to obtain the matching degree between the text field and the search question and the corresponding target matching degree, etc.
[0038] The text field can be a core field that is given a higher weight (boost) in the indexing and search process, which can also be referred to as a key text field. That is, the text field is mainly used to improve the ranking priority of the content of a specific field in the search result, and is intended to more accurately reflect the relevance between the user query intention and the document content through field-level weight control. For example, the text field can include title, name, product name, brand, author, label, keyword, abstract, and introduction, etc. For example, the title and abstract of text A can be used as text fields.
[0039] In the embodiments of the present disclosure, the text field can be a key field selected by the developer in the attribute set, which can be title, product name, film name, etc.
[0040] As known from the above introduction, in addition to the text field, the candidate set can also include a plurality of other fields, wherein the plurality of other fields can include numerical fields, category fields, time fields, geographic location fields, behavior characteristic fields, statistical characteristic fields, and custom fields, etc. The specific contents of the other fields will not be described here. The sorting method proposed in the present disclosure is mainly for the text field in the candidate set, and the processing of the other fields is similar to related technologies.
[0041] As an example, the search question is "meteor", and the plurality of text fields included in the corresponding obtained candidate set can be "meteor", "meteor, author: someone", and "a news article written by someone", etc.
[0042] It should be noted that when the to-be-searched question is received, the embodiment of the present disclosure can also perform intent recognition on the to-be-searched question to determine whether the scenario to which the to-be-searched question belongs is a specified scenario, and if it belongs to the specified scenario, the embodiment of the present disclosure can perform subsequent calculation of the matching type and the matching degree (fitting degree), wherein the specified scenario refers to a scenario in which the matching degree between the question and the answer is required to be relatively high, for example, the specified scenario can be a scenario in which the matching degree between the question and the answer exceeds a preset matching degree. Conversely, if the to-be-searched question does not belong to the target scenario, the embodiment of the present disclosure can not perform subsequent calculation of the matching type and the matching degree.
[0043] Optionally, if it is determined that the to-be-searched question does not belong to the target scenario, the embodiment of the present disclosure can also reduce the influence of the subsequent TBA feature (Tri-Band Affinity feature, three-frequency fitting factor feature), for example, for the scenarios of courses, comparisons and Q&A, the embodiment of the present disclosure can reduce the target matching degree.
[0044] In step S220, the matching type between the text field and the to-be-searched question is determined.
[0045] As an optional manner, after obtaining the candidate set including a plurality of text fields corresponding to the to-be-searched question, the embodiment of the present disclosure can traverse each text field in the candidate set to obtain the matching type between each text field and the to-be-searched question, and then obtain the target matching degree therebetween, so as to realize the sorting of each text field.
[0046] That is, the embodiment of the present disclosure can obtain the matching type between the text field and the to-be-searched question, and here, if the matching types are different, the ways of calculating the target matching degree corresponding thereto are also different. For example, the matching type can include an exact matching type (Exact), a substring matching type (SubString) and a bag-of-words matching type (BoW).
[0047] The matching type between the text field and the to-be-searched question can be obtained by comparing the string of the text field with the string of the to-be-searched question, so before determining the matching type between the text field and the to-be-searched question, the embodiment of the present disclosure can first perform a token operation on the text field and the to-be-searched question, respectively.
[0048] For example, for CJK, tokenization can be done on a per-character basis; for Latin text, tokenization can be done on a per-word basis, with the tokenization process allowing '-', '_', or '·' as intra-word connectors. Additionally, embodiments of the present disclosure can perform a deduplication operation on both sides of the token. Optionally, embodiments of the present disclosure can also utilize character n-gram BoW, or 2-3gram replacement tokens, to enhance the robustness of tokenization.
[0049] In embodiments of the present disclosure, the exact match type can also be referred to as the absolute match type, which refers to a type in which two strings are equal. In the process of determining the match type between the text field and the problem to be searched, embodiments of the present disclosure can determine whether the text field and the problem to be searched are equal, and if it is determined that the text field and the problem to be searched are equal, it is determined that the match type between the text field and the problem to be searched is the exact match type (Exact).
[0050] For example, the problem to be searched is a string Q, and the text field is a string T. When it is determined that T == Q (two strings are equal), it is determined that the match type between the text field and the problem to be searched is the exact match type (Exact). For example, the problem to be searched Q is "Artificial Intelligence", and the first text field T1 is the title of document 1, which is "Artificial Intelligence". By comparison, it can be seen that the first text field T1 is completely the same as the problem to be searched Q, so the match type between the first text field T1 and the problem to be searched Q is the exact match type (Exact).
[0051] Optionally, the substring match type (SubString) can also be referred to as the complete substring type, which refers to a type in which the field preferentially contains the continuous substring of the query (problem to be searched). In the process of determining the match type between the text field and the problem to be searched, if it is determined that the text field and the problem to be searched are not equal, embodiments of the present disclosure can determine whether the text field continuously includes the problem to be searched, and if it is determined that the text field continuously includes the problem to be searched, it is determined that the match type between the text field and the problem to be searched is the substring match type (SubString).
[0052] For example, the search question is a string Q, and the text field is a string T. When it is determined that T continuously contains Q and T≠Q, it is determined that the matching type between the text field and the search question is a substring matching type (SubString). For example, the search question Q is "artificial intelligence", and the second text field T2 is the title of document 2, which is "development of artificial intelligence". It can be determined that the second text field T2 contains the search question Q (continuously appears), and the second text field T2 "development of artificial intelligence" is not equal to the search question Q "artificial intelligence". Therefore, the matching type between the second text field T2 and the search question Q is a substring matching type (SubString).
[0053] Optionally, the dispersed matching type (BoW) can also be referred to as a bag-of-words matching type, which refers to a type that cannot be matched based on the order of tokens. In the process of determining the matching type between the text field and the search question, if it is determined that the text field does not continuously contain the search question, it is determined that the matching type between the text field and the search question is a dispersed matching type (BoW). In other words, when it is determined that the matching type between the text field and the search question does not satisfy the full matching type or the substring matching type, it is determined that the matching type between the text field and the search question is a dispersed matching type.
[0054] For example, the search question is a string Q, and the text field is a string T. When it is determined that T continuously contains Q and T≠Q, it is determined that the matching type between the text field and the search question is a substring matching type (SubString). For example, the search question Q is "artificial intelligence", and the third text field T3 is the title of document 3, which is "intelligent neural network overview". It can be determined that the third text field T3 contains part of the words of the search question Q "intelligent", and therefore the matching type between the third text field T3 and the search question Q is a dispersed matching type (BoW).
[0055] It should be noted that, before determining the matching type between the text field and the search question, the embodiments of the present disclosure can perform a truncation (Truncate / Trunc) operation on the text field and the search question. For example, the embodiments of the present disclosure can use Unicode code point truncation. By using Unicode safe truncation, the latency can be reduced and the determination can be stabilized without modifying the original field.
[0056] For example, the formula for truncating the search question is Q' = Trunc(Q, N), which means truncating the search question Q to at most N characters / code points to obtain a new search question Q'; the formula for truncating the text field is T' = Trunc(T, N), which means truncating the text field T to at most N characters / code points to obtain a new text field T'.
[0057] As an example, N is 10, the number of characters of the search question is 6, and the number of characters of the text field is 8. In this case, neither the search question nor the text field needs to be truncated, and the value of the truncated identifier can be false.
[0058] As another example, N is 10, and the search question is "Application and development trend of artificial intelligence in the medical field", which has 17 characters. By comparison, it is greater than N, which meets the truncation condition, and can be truncated to obtain a search question Q' of "Artificial intelligence in the medical field". In addition, the text field is "Research paper report on application of artificial intelligence technology in medical diagnosis 2024", which has 22 characters. By the same token, it meets the truncation condition and can be truncated to obtain a text field T' of "Artificial intelligence technology in medical diagnosis".
[0059] In summary, after obtaining the search question and the plurality of text fields, the embodiments of the present disclosure can first preprocess the two to perform a truncation operation on the strings that meet the truncation condition, and in the process of performing the truncation operation, the first N code points (characters) can be extracted as the basis for subsequent operations. Preferably, N can be between a specified range (64, 256), such as N can be 100.
[0060] The embodiments of the present disclosure can ensure that matching only the key field is greater than matching only the non-key field, and within the key field, complete matching is greater than continuous substring, and contact substring is greater than scattered matching, while retaining the overall relevance dominance, avoiding false positives that fields fit but the overall is irrelevant.
[0061] In step S230, the target matching degree between the text field and the search question is obtained based on the matching type.
[0062] As known from the above description, the matching type between the text field and the search question is not the same, and the way of obtaining the target matching degree is also not the same. That is, the embodiments of the present disclosure can obtain the target matching degree between the text field and the search question based on the matching type.
[0063] As an optional manner, in a case that the matching type between the text field and the to-be-searched question is determined as the exact matching type, the embodiment of the present disclosure can obtain a first matching degree (W Exact ) corresponding to the exact matching type, and take the first matching degree (W Exact ) as the target matching degree. The first matching degree (W Exact ) can be a default matching degree corresponding to the exact matching type, which is pre-set, for example, W Exact ≈3.0.
[0064] In addition, the first matching degree can also be dynamically adjusted according to the search demand of the user. For example, when the demand of the user is accurate matching, the embodiment of the present disclosure can increase the first matching degree, that is, the first matching degree can be greater than the default matching degree, for example, the first matching degree can be 4 or 5. Alternatively, when the demand of the user is not accurate matching and the requirement for accurate matching is not high, the embodiment of the present disclosure can decrease the first matching degree, for example, the first matching degree can be 2.8.
[0065] In summary, in a case that the to-be-searched question is equal to the text field, the target matching degree w =W Exact , that is, the matching type match_kind=Exact at this time is the exact matching type.
[0066] As an example, the to-be-searched question is “liuxing”, the text field is “liuxing”, the target matching degree w =W Exact =3.
[0067] As another optional manner, in a case that the matching type is determined as the substring matching type, the embodiment of the present disclosure can obtain the substring coverage degree between the to-be-searched question and the text field, and can also obtain a second matching degree and a third matching degree corresponding to the substring matching type. The second matching degree can be less than the third matching degree, and the third matching degree can be less than the first matching degree corresponding to the exact matching type. Based on this, the embodiment of the present disclosure can determine the target matching degree according to the substring coverage degree, the second matching degree and the third matching degree.
[0068] Here, the second matching degree can be a minimum matching degree (W sub_min ) corresponding to the substring matching type, and the third matching degree can be a maximum matching degree (W sub_max ) corresponding to the substring matching type. For example, the minimum matching degree W sub_min corresponding to the substring matching type can be 2, and the maximum matching degree W sub_max corresponding to the substring matching type can be 2.6, which is less than the first matching degree “3” corresponding to the exact matching type.
[0069] In addition, the sub-string coverage can be determined based on a length relationship between the search question and the text field, which can be calculated by the following formula: cov_sub = len(Q) / len(T); wherein cov_sub is the sub-string coverage, len(Q) is the string length of the search question, and len(T) is the string length of the text field. For example, the search question Q is "artificial intelligence", and the text field T is "artificial intelligence technology development report", then cov_sub = 4 / 10.
[0070] In addition, the embodiment of the present disclosure can determine the target matching degree by the following formula: w sub = W sub_min + (W sub_max - W sub_min ) × (cov_sub) α ; wherein w sub is the target matching degree corresponding to the sub-string matching type when the matching type between the search question and the text field is the sub-string matching type, and a is a constant.
[0071] In summary, when it is determined that the text field continuously includes the search question, the embodiment of the present disclosure can calculate the sub-string coverage cov_sub, and map the target matching degree w ∈ [W sub_min , W sub_max ], that is, the matching type match_kind = Sub (sub-string matching type) at this time.
[0072] As an example, the search question is "meteor", and the text field is "meteor author: someone", the range of cov_sub can be 0.33-0.5, and the range of the target matching degree w can be 2.2-2.35.
[0073] As another optional way, in the case of determining that the matching type is the scattered matching type, the embodiment of the present disclosure can obtain the bag-of-words coverage between the search question and the text field, and can also obtain the fourth matching degree and the fifth matching degree corresponding to the scattered matching type. Wherein the fourth matching degree can be less than the fifth matching degree, and the fifth matching degree can be less than the second matching degree corresponding to the above-mentioned sub-string matching type. Based on this, the embodiment of the present disclosure can comprehensively determine the target matching degree according to the bag-of-words coverage, the fourth matching degree and the fifth matching degree.
[0074] Here, the fourth matching degree can be the minimum matching degree (W bow_min ) corresponding to the scattered matching type, and the fifth matching degree can be the maximum matching degree (W bow_max ) corresponding to the scattered matching type. For example, the minimum matching degree W bow_minmay be 1, and the maximum matching degree W corresponding to the dispersion matching type bow_max may be 1.7, and the maximum matching degree "1.7" is less than the minimum matching degree "2" corresponding to the substring matching type.
[0075] In addition, the bag-of-words coverage can be determined based on the matching between the problem to be searched and the text field, which can be calculated by the following formula: Cov_bow= |TokQ∩TokT| / |TokQ|; Wherein, Cov_bow is the bag-of-words coverage, TokQ∩TokT is the number of vocabulary in the text field that exists in the problem to be searched, and TokQ is the number of vocabulary in the problem to be searched. For example, the problem to be searched Q is "smartphone price", TokQ=3, and the text field T is "phone shooting effect display", it can be seen that the same vocabulary between the problem to be searched Q and the text field is "phone", that is, |TokQ∩TokT|=1, so cov_bow=1 / 3.
[0076] In addition, the target matching degree can be determined by the following formula according to the embodiment of the disclosure: W bow = W bow_min + (W bow_max - W bow_min )×(cov_bow) β ; Wherein, w bow is the target matching degree corresponding to the dispersion matching type when the matching type between the problem to be searched and the text field is the dispersion matching type, and β is a constant.
[0077] The mapping function family (cov_sub) α , (cov_bow) β and the like can be replaced by any one of a piecewise linear function, a monotone spline function or a learnable function with a monotone constraint. These functions are monotonically non-decreasing with respect to coverage, and the output is limited to the corresponding frequency band interval, so the strict frequency band upper bound order and the intra-class monotonicity can be guaranteed. The mapping model with strict order relationship and intra-class monotonicity can solve the problems of experience weight jump and unprovable.
[0078] In summary, when it is determined that the text field only includes part of the vocabulary of the problem to be searched, that is, the text field includes all or part of the vocabulary of the problem to be searched, the bag-of-words coverage cov_bow can be calculated, and the target matching degree w∈[W bow_min , W bow_max ] is mapped, and the matching type match_kind=BoW (bag-of-words matching type) at this time. Before this, the token deduplication operation can also be performed to avoid repeated calculation.
[0079] As an example, the problem to be searched is "meteor", the text field only contains part of the vocabulary of the problem to be searched, the range of cov_bow can be 0.33-0.66, and the range of the target matching degree w can be 1.23-1.46.
[0080] In the embodiments of the present disclosure, the first matching degree W Exact The third matching degree W sub_max The fifth matching degree W bow_max That is, the first matching degree can be greater than the third matching degree, and the third matching degree can be greater than the fifth matching degree.
[0081] As another optional way, after obtaining the target matching degree, the embodiments of the present disclosure can also select a target text field that meets the specified condition from at least one text field, and on this basis, perform a length suppression operation on the target matching degree corresponding to the target text field to obtain a suppressed matching degree. In other words, before performing matching degree calculation on the text field, the embodiments of the present disclosure can determine whether the text field meets the specified condition, and if it is determined that the text field meets the specified condition, it can be used as a target text field. Subsequently, the embodiments of the present disclosure can perform a suppression operation on the target matching degree corresponding to the target text field to obtain a suppressed matching degree.
[0082] Here, the specified condition can include at least one of the following: the matching type between the text field and the problem to be searched is the exact matching type; and the character length of the text field is less than or equal to the length threshold.
[0083] As an example, when it is detected that the matching type between the text field and the problem to be searched is the exact matching type (Exact), the embodiments of the present disclosure can use the text field as a target text field, and subsequently, when the target matching degree corresponding to the target text field is obtained, the embodiments of the present disclosure can perform a length suppression operation on the target matching degree.
[0084] As another example, when it is detected that the character length of the text field is less than the length threshold, the embodiments of the present disclosure can use the text field as a target text field, and subsequently, when the target matching degree corresponding to the target text field is obtained, the embodiments of the present disclosure can perform a length suppression operation on the target matching degree.
[0085] As another example, when it is detected that the matching type between the text field and the to-be-searched question is an exact matching type (Exact), and the character length of the text field is less than a length threshold, the text field can be taken as a target text field, and subsequently, when a target matching degree corresponding to the target text field is obtained, the length suppression operation can be performed on the target matching degree.
[0086] In some embodiments, the length suppression operation can include a first-order suppression operation and / or a second-order suppression operation. The first-order suppression operation can be used to suppress the target matching degree based on the length threshold to obtain a suppressed matching degree. The second-order suppression operation can be used to obtain the suppressed matching degree according to a relationship between an obtained soft upper limit value and the target matching degree, wherein the soft upper limit value can be used to impose an adjustable upper limit on the target matching degree to dynamically adjust it according to risk.
[0087] The first-order suppression operation can also be referred to as a linear suppression operation, and the formula corresponding to the first-order suppression operation is: w' = w × P_len(L t ), wherein P_len(L t ) can be calculated by the following formula: P_len(L t ) = 1 γ × (1 L t / L0)^η, wherein P_len(L t ) is a length-aware suppression, L t is the text length after performing the truncation operation on the text field, L0 is the length threshold, which can be greater than 6 and less than 20, γ is the intensity, which can be greater than 0 and less than 0.3, and the index η is preferably 1-2. Based on the above formula, the shorter the length of the text field, the more suppression, and vice versa, the longer the length of the text field, the less suppression.
[0088] Optionally, the second-order suppression operation can be referred to as a length-related upper limit suppression operation, which sets an upper limit for the text field with a length less than a preset threshold to prevent the target matching degree corresponding to the text field from being too large. As known from the above introduction, the second-order suppression operation is used to obtain the suppressed matching degree according to a relationship between an obtained soft upper limit value and the target matching degree. The soft upper limit value W cap(Lt) = W cap_hi δ × exp( L t / τ), wherein δ and τ are parameters for controlling the lower limit and the rising speed, respectively, and W cap_hiis an upper limit when the long title is present, which can be a preset value. Exemplarily, W cap(Lt) may be greater than or equal to 3.0 and less than 3.5, such as W cap(Lt) may be 3.
[0089] Exemplarily, when the matching type is the full matching type, W cap_hi = W Exact (the first matching degree), δ = 0.6-0.9, τ = 4-6, when L t = 1, the upper limit range can be 2.1-2.4; when L t = 6, the upper limit range can be 2.7-2.9; when L t approaches L0, can approach 3.0.
[0090] In the process of obtaining the suppression matching degree according to the relationship between the soft upper limit value and the target matching degree, the embodiment of the disclosure can determine whether the target matching degree is less than the soft upper limit value calculated by the above-mentioned manner, if yes, the length suppression operation is not performed. Conversely, if the soft upper limit value is less than the target matching degree, the soft upper limit value can be taken as the suppression matching degree, and the corresponding calculation formula can be: w' = min(w, W cap(Lt ).
[0091] Through the second-order suppression operation, the embodiment of the disclosure can perform suppression processing on extremely short, abnormally popular or can be used by operation / garbage words to avoid amplification to an excessively high score.
[0092] The embodiment of the disclosure can determine to use the first-order suppression operation and / or the second-order suppression operation to suppress the target matching degree according to actual conditions. For example, when the online initial access sorting model is used, the embodiment of the disclosure can use the first-order suppression operation on the text field that meets the specified condition. For example, for the text field that is extremely short and peak sensitive, the embodiment of the disclosure can use the second-order suppression operation. For example, for the target text field that meets the specified condition, the embodiment of the disclosure can first perform the first-order suppression operation and then perform the second-order suppression operation, such as for high-risk business, the embodiment of the disclosure can first perform the first-order smoothing and then perform the second-order peak cutting, so that the target matching degree can be reduced by an average of one, and the peak can be prevented. How to perform length suppression on the text field that meets the specified condition is not limited here, and can be selected according to actual conditions.
[0093] In summary, in order to ensure that the very short title is over-amplified in the exact match / substring match, resulting in being used by the operation or garbage content, the embodiment of the disclosure can perform a short text penalty operation (length threshold operation). In other words, after determining the target matching degree, the embodiment of the disclosure can apply inhibition to the short text. Based on the above introduction, it is known that the length inhibition operation is only enabled in a specified scenario, such as when the match kind is the exact match type (Exact), or when the length of the text field |Lt| ≤ length threshold L0, the length inhibition operation is performed. Conversely, if any of the above conditions is not met, the length inhibition operation is not performed.
[0094] It should be noted that before determining whether the text field meets the specified condition, the embodiment of the disclosure can determine whether the text field is a target sample, and if it is a target sample, the length inhibition operation is not performed on it, wherein the target sample can be a prior positive sample. For example, when the text field is a brand or an official short name, even if the character length thereof meets the specified condition, the embodiment of the disclosure does not perform the inhibition operation thereon.
[0095] By performing the length inhibition operation, the embodiment of the disclosure can provide an optional length-aware inhibition strategy for very short titles / keyword fields to inhibit manipulation and noise without affecting regular samples.
[0096] In addition, before determining the match type between the text field and the to-be-searched question, the embodiment of the disclosure can also determine whether the text field is a sensitive field, and if the text field is a sensitive field, the match type acquisition operation is not performed on the text field, and the target matching degree acquisition operation is not performed. For example, when the hit noise word / sensitive word / operation volume word / known cheating word, the embodiment of the disclosure can directly turn off the entire acquisition process of the target matching degree, or weaken the influence of the target matching degree. For example, the target matching degree is set to 1, or only the weak gain brought by BoW is retained.
[0097] It should be noted that the determination process of the target matching degree can be implemented by a first sorting layer of a search system, that is, the first sorting layer can receive a candidate set output by a recall layer, and determine the match type between each text field in the candidate set and the to-be-searched question, and acquire the target matching degree between the text field and the to-be-searched question based on the match type. On this basis, the first sorting layer can sort at least one text field according to the target matching degree to obtain a sorting result.
[0098] In addition, after obtaining the target matching degree, the embodiment of the disclosure can also calibrate the target matching degree, such as calibrating the target matching degree by Platt / Isotonic to obtain a calibrated matching degree w cal, so as to avoid dimensional conflicts with other features of the model. The embodiment of the disclosure can transmit the calibrated matching degree to the downstream layer. In addition, when training the ranking model or the deep model, the embodiment of the disclosure can impose a monotonic non-decreasing constraint on the calibrated matching degree to ensure that the stronger is not reduced.
[0099] In step S240, at least one text field is ranked according to the target matching degree to obtain a ranking result.
[0100] As an optional way, after obtaining the target matching degree, the embodiment of the disclosure can rank at least one text field based on the target matching degree to obtain a ranking result. As known from the above introduction, the embodiment of the disclosure can be applied to the search system, which can include a recall layer, a plurality of ranking layers, and a rearrangement layer, etc. The calculation of the above target matching degree can be realized by the first ranking layer in the plurality of ranking layers.
[0101] In the process of ranking at least one text field based on the target matching degree, the embodiment of the disclosure can first obtain a base score corresponding to the text field, and then obtain a target ranking score based on the obtained target matching degree and the base score, and then rank at least one text field according to the target ranking score to obtain a ranking result.
[0102] The base score can be a ranking score obtained by ranking the text field by the ranking model, which can be the output of the previous level, such as the ranking score transmitted by the recall layer to the first ranking layer. The target ranking score can be obtained based on the following first formula: final_score = (max(base_score, 0) + ε) × w; Wherein, base_score is the base score, w is the target matching degree, and ε is a constant.
[0103] Based on the above formula, the embodiment of the disclosure can map the three types of matching signals (Exact / Substring / BoW) of the search question Q and the text field T to the weight w ≥ 1 through the continuous function, and then fuse them by multiplication in the ranking layer. Through the above multiplication fusion and non-negative protection in the ranking layer, the embodiment of the disclosure can avoid negative score amplification and ranking reversal.
[0104] Alternatively, the target ranking score can also be obtained based on the following second formula: final_score = base_score + × λ×(w 1); wherein λ is a constant. It is not specifically limited here which formula is used to obtain the target ranking score, and can be selected according to actual conditions.
[0105] It should be noted that after obtaining the target matching degree, the embodiment of the present disclosure can pass the target matching degree to the downstream layer, for example, the first ranking layer can transmit the obtained target matching degree to the second ranking layer or to the rearrangement layer.
[0106] As another optional way, the embodiment of the present disclosure can generate a target feature based on the first ranking layer, which can also be referred to as a three-frequency matching factor (TBA) feature, wherein the three-frequency matching factor involves three types of matching signals, including exact matching (Exact), substring matching (SubString) and dispersion matching (BoW).
[0107] The target feature can include a matching type (match_kind) and a target matching degree (weight_w). On this basis, the target feature is sent to at least one second ranking layer, wherein the second ranking layer (other ranking layer) can be used to perform a ranking operation based on the target feature. In other words, after obtaining the target matching degree, the embodiment of the present disclosure can generate a target feature from the target matching degree and other information. Here, the second ranking layer does not specifically refer to a specific ranking layer, which can be any other ranking layer after the first ranking layer.
[0108] Optionally, in addition to the matching type and the target matching degree, the target feature can also include a substring coverage, a bag-of-words coverage, a first character length, a second character length and a truncation operation identifier, wherein the first character length is the length of the character after performing the truncation operation on the search problem; the second character length is the length of the character after performing the truncation operation on the text field.
[0109] In order to better understand the target feature (TBAFeature), the embodiment of the present disclosure gives the following Table 1: Table 1
[0110] The weight_w in Table 1 can be a target matching degree, which can be a character number greater than or equal to 1, the match_kind can be a matching type, which can include three matching types, namely Exact, Sub and BoW, the cov_sub can be a substring coverage, the cov_bow can be a bag of words coverage, the len_q_trunc can be a first character length, the len_t_trunc can be a second character length, and the truncated can be a truncation operation identifier, the value of which is true when the truncation operation is performed on the string, and the value of which is false when the truncation operation is performed on the string.
[0111] After obtaining the target feature (TBA Feature), the embodiments of the present disclosure can also pass the target feature to other ranking layers. The other ranking layers can take the target feature as input and perform a ranking operation based on the target feature. By introducing the target feature, the other ranking layers can more accurately implement the ranking operation, i.e., obtain a more accurate ranking result. For example, the first ranking layer can transmit the target feature to at least one second ranking layer. Here, the second ranking layer can be another ranking layer other than the first ranking layer. In addition, after generating the target feature, the embodiments of the present disclosure can write it into metadata to facilitate its transmission to other ranking layers.
[0112] It should be noted that the embodiments of the present disclosure can transmit the target feature (TBA Feature) to multiple other ranking layers through a feature bus (Feature Bus). Meanwhile, the embodiments of the present disclosure can cache the target feature. Exemplarily, the cached information can include tenant_id (tenant identifier), dataset_id (dataset identifier), doc_id (document ID), query_hash (query hash), and priority_text_hash (priority text hash).
[0113] As an example, the search problem to be searched is "How to raise green plants", and the corresponding obtained ranking result can be as shown in Table 2. Figure 3 The more matched the title name is with the search problem to be searched, the greater the corresponding target matching degree is, and the higher the final ranking score is, i.e., the result presented to the user is earlier.
[0114] In summary, after obtaining the target matching degree or the target feature, the embodiment of the disclosure can transmit the target matching degree and / or the target feature to multiple ranking layers to realize calibration and individual constraint of the target matching degree. That is, the downstream ranking layer can directly use the target matching degree multiplication fusion, that is, each downstream ranking layer can fuse the target matching degree transmitted by the first ranking layer with the basic ranking score after obtaining the basic ranking score to obtain a more accurate ranking score. Alternatively, other downstream ranking layers can also input the target feature (TBA Feature) into a ranking learning model (LTR model), so that calibration (Platt / Isotonic) and monotonic constraint can be performed, thereby ensuring online consistency. In addition, the embodiment of the disclosure can also avoid repeated calculation through caching.
[0115] It should be noted that the target feature can be transmitted not only to other ranking layers but also to a rearrangement layer. For example, the embodiment of the disclosure can input the target feature (TBA Feature) as an input of a rearrangement model (rearrangement layer), which can be input in original or discretization / embedding form. In addition, the offline sample can use the same Trunc / token rule and TBA Feature generation process as the online sample, and the embodiment of the disclosure can freeze and version the feature dictionary to avoid drift.
[0116] In addition, in the process of transmitting the target feature (TBA Feature), the embodiment of the disclosure can transmit the target feature in the feature bus through candidate metadata with the request, and the rearrangement layer (rearrangement stage) can directly read it without repeated calculation. In this process, the P95 ranking delay is basically unchanged, and cross-layer reuse can avoid repeated calculation, that is, shared matching features, thereby reducing the ranking jitter caused by inter-layer logical differences, and better interpretability.
[0117] It is worth noting that if an abnormality of the rearrangement layer is detected, the embodiment of the disclosure can degrade to use the upstream target matching degree w multiplication. For example, when the calibration matching degree w cal is unavailable, it can be degraded to the target matching degree.
[0118] It can be seen that, without changing the rearrangement layer, the embodiment of the disclosure can make the rearrangement layer guarantee the “field matching degree” and keep consistent with the upstream ranking layer by providing the standardized feature (target feature) of the TBA as the model input. In addition, this method can reduce the access cost.
[0119] It should be noted that, for the same search result, the embodiment of the present disclosure can perform a multi-field group priority operation, such as taking the maximum value or average value within a title+brand group, or aggregating according to a weight. That is, when the same result corresponds to multiple text fields, the embodiment of the present disclosure can weight or average these text fields. For example, the search result "Article A" corresponding to the search question corresponds to three text fields, which include the name, abstract and author respectively. When the target matching degrees of the three text fields are obtained, the embodiment of the present disclosure can perform a combination processing on the three target matching degrees, such as taking the maximum value or weighted summation, to obtain the target matching degree corresponding to the search result.
[0120] The embodiment of the present disclosure can improve the stability and consistency of entity query sorting to some extent by modeling and weighting the matching degrees of the text fields and the search question (query text) using the sorting layer, and using the target matching degree as a standardized feature (target feature) in multiple sorting layers. For example, the stability and consistency of the title, product name and film name can be improved, and multi-layer reuse can reduce repeated implementation and inter-layer jitter. In addition, the embodiment of the present disclosure clearly defines the training / online alignment and calibration strategy, which can ensure the consistency of the whole link.
[0121] Sorting multiple text fields based on target matching degrees can improve the hit rate of entity queries, not only ensuring continuous substring steady-state improvement, but also allowing dispersed hits to be moderately improved. Based on this, the first-screen CTR (Click-Through Rate), play rate, complete play rate, ATC (Add To Cart), conversion rate and the like can be improved. In addition, the embodiment of the present disclosure supports attribute set configuration and single priority field selection, which facilitates covering news / e-commerce / audio / video and other industries.
[0122] Based on the same inventive concept, the present disclosure also provides a sorting device, Figure 4 Fig. 4 is a block diagram of a sorting device 400 according to an example embodiment. Figure 4 As shown in the figure, the sorting device 400 can include a first acquisition module 410, a determination module 420, a second acquisition module 430 and a sorting module 440.
[0123] The first acquisition module 410 is configured to receive a search question, acquire a candidate set corresponding to the search question, and the candidate set includes at least one text field; The determination module 420 is configured to determine the matching type between the text field and the search question, and the matching type includes a complete matching type, a substring matching type and a dispersed matching type; The second obtaining module 430 is configured to obtain a target matching degree between the text field and the question to be searched based on the matching type. The sorting module 440 is configured to sort the at least one text field according to the target matching degree to obtain a sorting result.
[0124] In some embodiments, the second obtaining module 430 can also be configured to, when the matching type is the full matching type, obtain a first matching degree corresponding to the full matching type, and take the first matching degree as the target matching degree.
[0125] In some embodiments, the second obtaining module 430 can also be configured to, when the matching type is the substring matching type, obtain a substring coverage between the question to be searched and the text field, and obtain a second matching degree and a third matching degree corresponding to the substring matching type, the second matching degree being less than the third matching degree, and the third matching degree being less than the first matching degree; and determine the target matching degree according to the substring coverage, the second matching degree, and the third matching degree.
[0126] In some embodiments, the second obtaining module 430 can also be configured to, when the matching type is the dispersion matching type, obtain a bag-of-words coverage between the question to be searched and the text field, and obtain a fourth matching degree and a fifth matching degree corresponding to the dispersion matching type, the fourth matching degree being less than the fifth matching degree, and the fifth matching degree being less than the second matching degree; and determine the target matching degree according to the bag-of-words coverage, the fourth matching degree, and the fifth matching degree.
[0127] In some embodiments, the device is applied to a search system, and the search system includes a first sorting layer and at least one second sorting layer. The sorting device further includes: A multiplexing module configured to generate a target feature based on the first sorting layer, the target feature including the matching type and the target matching degree; and send the target feature to at least one second sorting layer, the second sorting layer being configured to perform a sorting operation based on the target feature.
[0128] In some embodiments, the target feature further includes at least one of the following information: a substring coverage; a bag-of-words coverage; a first character length, the first character length being a length of characters after a truncation operation performed on the question to be searched; a second character length, the second character length being a length of characters after a truncation operation performed on the text field; a truncation operation identifier.
[0129] In some embodiments, the ranking module 440 can be further configured to obtain a basic ranking score corresponding to the text field; obtain a target ranking score based on the target matching degree and the basic ranking score; and rank the at least one text field according to the target ranking score to obtain the ranking result.
[0130] In some embodiments, the ranking apparatus further comprises: a suppression module configured to select a target text field satisfying a specified condition from the at least one text field; and perform a length suppression operation on the target matching degree corresponding to the target text field to obtain a suppressed matching degree.
[0131] In some embodiments, the specified condition comprises at least one of: the matching type between the text field and the problem to be searched is the complete matching type; the character length of the text field is less than or equal to a length threshold.
[0132] In some embodiments, the length suppression operation comprises at least one of: a first-order suppression operation for suppressing the target matching degree based on a length threshold to obtain the suppressed matching degree; a second-order suppression operation for obtaining the suppressed matching degree according to a relationship between an obtained soft upper limit value and the target matching degree.
[0133] In the case of receiving a problem to be searched, the embodiments of the present disclosure obtain a candidate set corresponding to the problem to be searched, which can include a plurality of text fields, determine a matching type between the text field and the problem to be searched on this basis, and obtain a target matching degree between the text field and the problem to be searched based on the matching type, wherein the matching type includes a complete matching type, a substring matching type and an assignment matching type, and finally rank the plurality of text fields in the candidate set based on the target matching degree to obtain a ranking result. Since the candidate set ranking introduces the target matching degree, and the target matching degree is determined by the matching type between the text field and the problem to be searched, the accuracy of the priority field ranking can be greatly guaranteed.
[0134] Reference will be made to the following description of embodiments of the application Figure 5FIG. 5 shows a structural diagram of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of embodiments of the present disclosure.
[0135] As shown in FIG. 5, the electronic device 500 can include a processing device 501, which can be a central processor, a graphic processor, or the like, that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504. Figure 5 Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or via a wire to exchange data. Although
[0136] The electronic device 500 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or possessed. More or fewer devices can alternatively be implemented or possessed. Figure 5
[0137] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0138] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0139] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0140] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device without being incorporated into the electronic device.
[0141] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: receive a search question, obtain a candidate set corresponding to the search question, the candidate set comprising at least one text field; determine a matching type between the text field and the search question, the matching type comprising a complete matching type, a substring matching type, and a scattered matching type; obtain a target matching degree between the text field and the search question based on the matching type; and sort the at least one text field according to the target matching degree to obtain a sorting result.
[0142] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++, as well as conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0143] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It will also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.
[0144] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself. For example, the second acquisition module can also be described as "obtaining a target matching degree between the text field and the problem to be searched based on the matching type".
[0145] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0148] Moreover, while operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that some operations can be performed in parallel or in any suitably enabled order. Similarly, while several specific implementation details are included herein, they should not be taken as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
Claims
1. A sorting method, characterized in that, The method includes: Receive a question to be searched, and obtain a candidate set corresponding to the question to be searched, wherein the candidate set includes at least one text field; Determine the matching type between the text field and the search question, where the matching type includes exact match, substring match, and scattered match. Based on the matching type, obtain the target matching degree between the text field and the question to be searched; The at least one text field is sorted according to the target matching degree to obtain a sorting result.
2. The sorting method according to claim 1, characterized in that, The step of obtaining the target match degree between the text field and the search question based on the matching type includes: When the matching type is the exact matching type, the first matching degree corresponding to the exact matching type is obtained, and the first matching degree is used as the target matching degree.
3. The sorting method according to claim 1, characterized in that, The step of obtaining the target match degree between the text field and the search question based on the matching type includes: When the matching type is the substring matching type, obtain the substring coverage between the question to be searched and the text field, and obtain the second matching degree and the third matching degree corresponding to the substring matching type, wherein the second matching degree is less than the third matching degree, and the third matching degree is less than the first matching degree; The target matching degree is determined based on the substring coverage, the second matching degree, and the third matching degree.
4. The sorting method according to claim 1, characterized in that, The step of obtaining the target match degree between the text field and the search question based on the matching type includes: When the matching type is the scattered matching type, obtain the bag-of-words coverage between the question to be searched and the text field, and obtain the fourth matching degree and the fifth matching degree corresponding to the scattered matching type, wherein the fourth matching degree is less than the fifth matching degree, and the fifth matching degree is less than the second matching degree; The target matching degree is determined based on the bag-of-words coverage, the fourth matching degree, and the fifth matching degree.
5. The sorting method according to claim 1, characterized in that, The method is applied to a search system, the search system including a first ranking layer and at least one second ranking layer, and the method further includes: Target features are generated based on the first ranking layer, and the target features include the matching type and the target matching degree; The target feature is sent to at least one second sorting layer, which performs a sorting operation based on the target feature.
6. The sorting method according to claim 5, characterized in that, The target feature also includes at least one of the following information: Substring coverage; Bag-of-words coverage; The first character length is the length of the characters after the truncation operation is performed on the question to be searched; The second character length is the length of the characters after the text field has been truncated. Truncation operation identifier.
7. The sorting method according to any one of claims 1 to 6, characterized in that, The step of sorting the at least one text field according to the target matching degree to obtain a sorting result includes: Get the base sort score corresponding to the text field; The target ranking score is obtained based on the target matching degree and the basic ranking score; The at least one text field is sorted according to the target sorting score to obtain the sorting result.
8. The sorting method according to any one of claims 1 to 6, characterized in that, The method further includes: Select a target text field that meets the specified conditions from the at least one text field; A length suppression operation is performed on the target matching degree corresponding to the target text field to obtain the suppressed matching degree.
9. The sorting method according to claim 8, characterized in that, The specified conditions include at least one of the following: The matching type between the text field and the question to be searched is the exact match type; The character length of the text field is less than or equal to the length threshold.
10. The sorting method according to claim 8, characterized in that, The length suppression operation includes at least one of the following: A first-order suppression operation is used to suppress the target matching degree based on a length threshold to obtain the suppressed matching degree; The second-order suppression operation is used to obtain the suppression matching degree based on the relationship between the obtained soft upper limit value and the target matching degree.
11. A sorting device, characterized in that, The device includes: The first acquisition module is configured to receive a question to be searched and acquire a candidate set corresponding to the question to be searched, wherein the candidate set includes at least one text field. The determination module is configured to determine the matching type between the text field and the search question, the matching type including exact match, substring match and scattered match. The second acquisition module is configured to acquire the target matching degree between the text field and the question to be searched based on the matching type. The sorting module is configured to sort the at least one text field according to the target matching degree to obtain a sorting result.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-10.
13. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and system for dynamically rankings images to be matched with content in response to a search query
CN107463591A
Query result matching degree calculation method and device
CN111221943A
Text matching method and device, electronic equipment and storage medium
CN111931477A
Semantic understanding method and device, equipment and storage medium
CN113569565A
Clue data mining method and device based on Flink and multimode matching and medium
CN119127972A