A Calculation Method for Query Title Intention Sorting Weight (QTI) in Document Search
The QTI calculation method addresses the challenge of accurately ranking search results by evaluating word coverage and relevance in document titles and bodies, improving search accuracy and user satisfaction.
Patent Information
- Application Number
- CN202210695685.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-06-20
AI Technical Summary
The prior art is difficult to accurately understand the user's multi-step search intention in document search, resulting in inaccurate search results.
The calculation method of the QTI title intent sorting weight QTI is used for document search, and the QTI value is calculated by marking important words, calculating hit coverage and matching score values, and combining with the BM25 algorithm, the QTI value is calculated for sorting.
Improve the accuracy of search results and enhance user search satisfaction.
Smart Images

Figure CN115238061B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document search, and particularly relates to a method for calculating the sorting weight QTI of the title intention of a document search Query. Background Art
[0002] The explosive growth of data has put forward more requirements for search. Among many requirements, fast search response and accurate results are the most basic and concerned requirements. In terms of speed, there are many underlying search engines that provide good solutions. Among them, the open-source Solr and Elasticsearch have made important contributions to the search field. In most cases, after handing over the requirement of speed to the search engine, another concern is how to better utilize the search engine to obtain more accurate results.
[0003] Generally, when the words in the content searched by the user are close together, they have a clear meaning. For example, "Huawei mobile phone", obviously, the two words together have a more clear meaning. If they are hit separately, it is easy to give problematic results. For example, "A mobile phone was found near the Huawei campus", in this case, the keyword coverage rate is 1, but the meaning expressed is very different from the original query. In the world of search, it is not necessarily the case that the closer the keywords searched by the user are together, the more likely they are the results the user wants. For example, in the field of document search, for "Pangu test process", what the user wants to search for is the test process of the Pangu product. Usually, the title of the document is "Pangu Development Specification", and then there is a chapter "Test Process" inside the document. At this time, the content searched by the user is very far apart in the results to be hit; in this case, the user's search is actually divided into two steps: First, the user expects to find documents related to "Pangu", and second, the user expects to find the "Test Process" in that document. This is the problem that QTI needs to solve.
[0004] Learning To Rank (LTR), as a ranking method that uses machine learning means to learn the user's search behavior and consider the similarity between the search query and the content to be searched from multiple dimensions, is essential when there is sufficient search log. When using the LTR technical means, on the one hand, some deep learning methods can be used to automatically mine feature information; secondly, manually mining good ranking features is an important means to improve the search effect; and in many internal enterprise search scenarios, the search log data is insufficient, and the way of manually mining features becomes even more important. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention provides a method for calculating the sorting weight QTI of the Query title intention in document search. QTI can be used as an important weight index for fine sorting the rough search results, or can be directly added as a feature to the LTR algorithm, serving as one of the dimensions for comprehensively considering search fine sorting to improve the accuracy of search results.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect of the present invention, there is provided a method for calculating the sorting weight QTI of the Query title intention in document search. The method is characterized by including the following steps:
[0008] Step 1: According to the search command, mark the important words to be recognized, and obtain the corresponding set of important words IW, IW = [w1, w2, w3,..., w M ;
[0009] Step 2: Select a document and calculate the hit coverage rate Cr of the important words in the article title Title(T); M is the total number of all important words in IW, 1 is the number of times the important words appear in the document title;
[0010] Step 3: Calculate the matching score value of the important words in the content of the document other than the title by means of the BM25 calculation method;
[0011] BM25(IW) = ∑BM25(w i ), i = 1, 2, 3,..., M;
[0012] Step 4: Calculate the hit coverage rate of the non-important words other than the important words in the search command in the entire document; 1 is the number of times the non-important words appear in the document, and N is the value obtained by subtracting M from the number of words in the total word set corresponding to the search command Query(Q);
[0013] Step 5: Calculate QTI;
[0014] QTI = C r (IW) × (BM25(IW) + S * Cr(EIW)); S is the basic BM25 score of the search command Query(Q) and the document in the retrieval process;
[0015] Step 6: Sort according to the calculated QTI score and return the sorting result.
[0016] Further, a model for identifying important words is trained on the extraction platform. The input of this model is the user's search command Query (Q), and the output is the IW set.
[0017] Furthermore, in step one, the important words are the names of various items, and IW is a set without duplicates.
[0018] Furthermore, in step two, for the same important word, if it appears in the title, it is marked as 1 time, and if it does not appear in the title, it is marked as 0 time.
[0019] Furthermore, in step four, for the same non-important word, if it appears in the document, it is marked as 1 time, and if it does not appear in the document, it is marked as 0 time.
[0020] On the other hand, the present invention also provides a device for calculating the sorting weight QTI of the Query title intention in document search, including:
[0021] A platform extraction module; the platform extraction module is used to extract important words related to the input command and form a set of important words;
[0022] A search engine; the search engine is used to query the input command;
[0023] A processor connected to the search engine and the platform extraction module.
[0024] On the other hand, the present invention also provides a chip, including a processor. The processor is used to call information from the platform extraction module and the search engine and execute a computer program, so that a device installed with the chip executes the method according to any one of claims 1-5.
[0025] On the other hand, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to perform the steps of the method according to any one of claims 1-5 above.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows: The QTI index calculated by the present invention can be directly used as the rough search result, as an important weight index for fine ranking, or can be directly used as a feature and added to the LTR algorithm as one of the dimensions for comprehensively considering search fine ranking, thereby improving the user's search satisfaction. Brief Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 It is a flowchart of the QTI calculation method of the present invention. Detailed implementation manners
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0030] As Figure 1 shown, this embodiment provides a calculation method for the Query Title Intent Sorting Weight (QTI) of document search. When a user searches for a Query, the Query is segmented and analyzed, and DSL statements are composed. Elasticsearch is used for querying, and the query results are sorted by QTI. The specific steps are as follows:
[0031] Step 1: According to the search command, mark the important words that may need to be recognized during the calculation process. This process can utilize the extraction function of the IDPS platform. Generally, such important words are the names of various items, such as product names, stock names, company names, and their corresponding abbreviations. After the model is trained, the input is the user's search command to obtain the Query (Q), and the corresponding set of important words IW (Important Word), IW = [w1, w2, w3,..., w M ; The IW set is a set without duplicates.
[0032] Step 2: Select a document and calculate the hit coverage ratio (Cr) of the important words in the article title Title (T);
[0033] M is the total number of all important words in IW, 1 is the number of times the important word appears in the document title; for the same important word, if it appears in the title, it is marked as 1 time, and if it does not appear in the title, it is marked as 0 time.
[0034] Step 3: Calculate the matching score value of such important words in the content of the document except the title by means of the BM25 calculation method; in the case where the important name identified in the Title is hit, the search is basically locked on the document, and a certain hit of this noun in the content is a bonus item.
[0035] BM25(IW) = ∑BM25(w i ), i = 1, 2, 3, …, M.
[0036] Step 4: Calculate the hit coverage rate of the non-important words other than the important words in the search command in the entire document; 1 is the number of times the non-important word (Except Important Word, EIW) appears in the document, and N is the value obtained by subtracting M from the number of words in the total word set corresponding to the search command Query (Q); for the same non-important word, if it appears in the document, it is marked as 1 time, and if it does not appear in the document, it is marked as 0 time.
[0037] Step 5: Calculate QTI;
[0038] QTI = C r (IW) × (BM25(IW) + S * Cr(EIW)); S is the basic BM25 score of the search command Query (Q) and the document in the retrieval process by the search engine.
[0039] Step 6: Sort according to the QTI calculation score and return the sorting result.
[0040] The present invention also provides a device for calculating the sorting weight QTI of the Query title intention of document search, including:
[0041] A platform extraction module; the platform extraction module is used to extract important words related to the input command and form a set of important words;
[0042] A search engine; the search engine is used to query the input command;
[0043] A processor connected to the search engine and the platform extraction module.
[0044] The present invention also provides a chip, including a processor, and the processor is used to call information from the platform extraction module and the search engine and execute a computer program, so that a device installed with the chip executes the above method.
[0045] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the steps of the above method.
[0046] Although the above embodiments have specifically described the present invention, those of ordinary skill in the art should understand that modifications or improvements can be made based on the disclosed content of the present invention without departing from the spirit and scope of the present invention, and these modifications and improvements are within the spirit and scope of the present invention.
Claims
1. A method for calculating the sorting weight QTI of the Query title intention in document search, characterized in that It includes the following steps: Step 1. According to the search command, mark the important words to be recognized, and obtain the corresponding set of important words IW, where IW = [w1, w2, w3, …, w M ; Step 2: Select a document and calculate the hit coverage rate C of important words in the article title Title(T) r ; Let \(M\) be the total number of all important words in \(IW\), and \(n\) be the number of times important words appear in the document title; Step 3: Calculate the matching score value of the important word in the content of the document except the title by means of the BM25 calculation method; BM25(IW) = ∑BM25(w i ), i = 1, 2, 3, …, M; Step 4: Calculate the hit coverage rate of the non-important words other than the important word in the search command in the entire document; is the number of occurrences of non-important words in the document, and N is the value obtained by subtracting M from the number of words in the total word set corresponding to the search command Query (Q); Step 5: Calculate QTI; QTI = C r (IW) × (BM25(IW) + S * C r (EIW)); S is the basic BM25 score of the search command Query (Q) and the document during the retrieval process by the search engine; Step 6: Sort according to the score calculated by QTI and return the sorting result.
2. The calculation method of the document search Query title intention sorting weight QTI according to claim 1, characterized in that, Train a model for identifying important words on the extraction platform. The input of this model is the search command Query (Q) of the user, and the output is the IW set.
3. The calculation method of the document search Query title intention ranking weight QTI according to claim 1, characterized in that, In Step 1, the important words are the names of each item, and IW is a set without duplicates.
4. The calculation method of the document search Query title intention ranking weight QTI according to claim 1, characterized in that In Step 2, for the same important word, if it appears in the title, it is marked as 1 time, and if it does not appear in the title, it is marked as 0 time.
5. The calculation method of the document search Query title intention sorting weight QTI according to claim 1, characterized in that, In Step 4, for the same non-important word, if it appears in the document, it is marked as 1 time, and if it does not appear in the document, it is marked as 0 time.
6. A computing device for calculating the sorting weight QTI of the title intention of a document search Query, characterized in that, Apply the calculation method of the document search Query title intention sorting weight QTI described in any one of claims 1-5. The device includes: A platform extraction module; the platform extraction module is used to extract important words related to the input command and form a set of important words; A search engine; the search engine is used to query the input command; A processor connected to the search engine and the platform extraction module.
7. A chip, characterized in that, It includes a processor, and the processor is used to call information from the platform extraction module and the search engine and execute a computer program, so that the device installed with the chip executes the method described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it realizes the steps of the method described in any one of the above claims 1-5.
Citation Information
Patent Citations
Short sentence retrieval method for internet search
CN108268663A
Intellectual property matching technology based on keyword extraction and word shift distance
CN111027306A