A Document Theme Relevance Model in Pseudo-Relevance Feedback
By integrating theme-based document relevance into PRF models, the reliability of feedback documents is assessed, improving retrieval performance through enhanced query expansion, notably in TopRM3 and TopRoc, addressing the unreliability issues of traditional PRF models.
Patent Information
- Application Number
- CN202210105565.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The existing pseudo-related feedback method fails to effectively consider the reliability of the feedback document when selecting candidate terms, resulting in vocabulary entries being added to the query representation, affecting the performance of the second pass retrieval.
By introducing the correlation between topic-based feedback documents in the language model, the reliability of the feedback documents is estimated, and the topic correlation information is integrated in the Rocchio model, the TopRM3 and TopRoc models are proposed, and the query representation is optimized using topic similarity and TF-IDF weights.
Improved retrieval performance of pseudo-correlation feedback, especially on RM3 and Rocchio models, exhibited higher robustness and improved MAP performance.
Smart Images

Figure CN114611490B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text retrieval, information retrieval or data mining, and particularly to a document topic relevance model in pseudo-relevance feedback. Background Art
[0002] Pseudo relevance, also known as blind relevance feedback, is an automatic local analysis method. It automates the manual operation part of relevance feedback, thereby improving retrieval performance. This method first performs a normal retrieval process, returns the most relevant documents to form an initial set, then assumes that the top k documents are relevant, and finally performs relevance feedback as usual based on this assumption.
[0003] Pseudo Relevance Feedback (PRF) through Query Expansion (QE) is generally considered a very effective method for achieving good performance in Information Retrieval (IR). Although the PRF model usually performs very well, it will fail in some cases. In the classic PRF model, such as the Rocchio model or the Relevance Model RM3, all the top k feedback documents are assumed to be equally relevant to the query. The weights of candidate terms are only based on their importance in the set. These models cannot determine the reliability of the candidate documents when selecting them. Generally, words with the same weight (such as the Term Frequency–Inverse Document Frequency (TF-IDF) index) in different feedback documents are considered equally reliable for QE. When some feedback documents cover different topics, many of which are irrelevant to the original query, the models adopting the classic PRF strategy (e.g., Rocchio and RM3) do not perform well. In this case, a large number of irrelevant terms related to irrelevant topics in the documents are also added to the new query representation, which has a negative impact on the retrieval performance of the second-pass retrieval.
[0004] Recently, researchers have started to use topic models for PRF to obtain feedback words for the most relevant topics. However, most of their methods select candidate terms from the top-k documents returned in the first-round retrieval without considering the reliability of these documents. Since the original queries are generally short and their topics are ambiguous, the current methods have significant drawbacks. To address this issue, Miao et al. proposed a probabilistic framework TopPRF by integrating "topic space" (TS) information into the Rocchio model, which estimates the reliability of feedback documents by considering the relevance between the top-3 documents and other documents. Summary of the Invention
[0005] Object of the Invention: From the background description, most studies have not determined the reliability of the top-k documents returned in the first-pass retrieval when selecting candidate terms. To solve the above problem, the present invention proposes to estimate the reliability of feedback documents by introducing the relevance between topic-based feedback documents in the language model. Different from the work proposed by Miao et al., the method of the present invention can be considered a general method and can be incorporated into any other PRF model. The relevance model of topic-based pseudo-relevance feedback proposed by the present invention was verified on 5 public datasets of TREC, and the experimental results show that the algorithm of the present invention has good performance.
[0006] Technical Solution:
[0007] The relevance model of the present invention is based on the classical pseudo-relevance feedback framework and realizes (pseudo) relevance feedback by improving the query representation method. In the relevance model RM1, the weight of a candidate term w is: P(w|R) ∝ ∑ D∈F P(w|D)·P(D|Q) (1)
[0008] where Q is a query, D is a document in the feedback document set F, P(w|D) is the document language model, and P(D|Q) is the query language model. The present invention adopts this framework but utilizes the topic-based document relevance P T (D|F) instead of estimating P(D|Q) from the document scores. The topic-based relevance model proposed by the present invention is as follows:
[0009] P T (w|R) ∝ ∑ D∈F P(w|D)·P T (D|F) (2)
[0010] Similarly, the variant RM3 of the relevance model is the original query model θ Q and the feedback language model θ FA linear combination is performed, and the corresponding formula is as follows:
[0011] θ Q =(1 - α)·θ Q +α·θ F (3)
[0012] RM3 not only achieves good performance in terms of precision and recall metrics, but also has more prominent robustness. The Dirichlet prior for smoothing the document language model
[21] is used therein. On this basis, fully considering the above strategies, the present invention uses the topic-based relevance model P T (w|R) to estimate the feedback language model θ F , and proposes the TopRM3 model.
[0013] Given a query Q, the top-k feedback document set F in the first-pass retrieval can be represented by their topic distributions P T (z|D) (D ∈ F) in the topic space. However, P T (z|D) cannot characterize the topic-based relevance directly measured between the query and the document. The reason is that the topic distribution representation of short queries is usually very sparse and rough. Related research work
[10] finds that the first s (s << k) documents in the feedback document set are most likely to be relevant to the query topic. Usually, the relevant documents for a specific query cover several topics. To maintain the balance between topic diversity and document relevance, in the present invention, the value of s is set to 3, and the topics in the first s documents are regarded as the query topics.
[0014] The present invention measures the similarity between the topic vectors of two documents through the cosine formula. For document D i and D j , the topic similarity is as follows:
[0015]
[0016] where z is the topic and P(z|D) is a topic distribution in D. When s = 3, the topic similarity TS(D) is calculated as follows:
[0017]
[0018] where the exponents of the first three documents are set to 1 because by default they are relevant to the query.
[0019] Next, the topic similarity TS(D) is converted into the topic-based relative relevance of the document P T (D|F). Since P T(D|F) is a distribution, and two normalization schemes can be adopted, namely the linear method and the Soft-max method, as follows:
[0020]
[0021]
[0022] Among them, formula (6) represents the linear method, and formula (7) represents the Soft-max method.
[0023] Based on the above two correlation representation methods, the present invention proposes a topic-based correlation model, which are respectively represented as TopRM3-L and TopRM3-S.
[0024] The present invention further integrates the topic-based correlation information into the Rocchio model, that is, the topic-based Rocchio model TopRoc is obtained. The specific description is as follows:
[0025] (1) All documents are ranked for a given query using a specific IR model. BM25 is used in the first pass of retrieval, and the top |F| ranked documents are determined as the pseudo-relevant set F.
[0026] (2) Each candidate term in the top |F| ranked documents is assigned an expansion weight. Generally, the expansion weight is the dot product of the weights provided by the weighting model and the topic-based document correlation. The present invention uses the TF-IDF model
[22] as the weighting model.
[0027] (3) The vector of query term weights is a linear combination of the initial query term weights and the expansion weights. The formula is as follows:
[0028] Q1 = α·Q0 + β·∑ D∈F r(D)·P T (D|F) (8)
[0029] In the formula, Q0 and Q1 represent the original query vector and the query vector generated after one iteration respectively, α and β are adjustment parameters that control the degree of dependence on the original query vector and the feedback information, r(D) is the TF-IDF weight vector of the feedback document D, F is the feedback document set of PRF, and P T (D|F) measures the degree of topic correlation of the feedback document D. In practice, α can always be fixed at 1, and only β is studied to obtain better performance. If P T (D|F) follows a uniform distribution, the TopRoc model is the original Roccho model. In addition, the previously defined P TTwo calculation methods of (D|F), and the corresponding models are respectively represented as TopRoc-L and TopRoc-S.
[0030] Advantages: The present invention has been verified on 5 publicly available TREC datasets and achieved good experimental results. The results in Table 1 show that RM3 is better than LM on all collections, and the Rocchio model is better than BM25. This indicates that RM3 and the Rocchio model are still very powerful basic models in IR research work. Compared with the Rocchio model, RM3 performs more stably when the number of feedback documents changes. In Table 1, the TopRM3 model proposed by the present invention is better than RM3, and the TopRoc model is better than the Rocchio model on all collections. In particular, TopRM3-L and TopRoc-L have obvious improvements compared with the RM3 and Rocchio models respectively. In most cases, TopRoc-L obtains the best performance in terms of MAP. In Table 2, the present invention also compares the proposed TopRoc with TS-COS. TS-COS is the most effective variant of TopPRF and is considered to be one of the most effective and state-of-the-art PRF models. For a fair comparison, the improvement percentages of them compared with the corresponding Rocchio model are used here. The results are shown in Table 2. From Table 2, the proposed TopRoc-L has a greater improvement in most cases, indicating that the theme-based relevance model proposed by the present invention is more effective.
[0031] In addition, the sensitivity of the proposed TopRM3-L of the present invention to the number of feedback files and the feedback interpolation coefficient α in terms of MAP is studied. The results are as Figure 1 and Figure 2 shown. In Figure 1 , when the number of feedback documents changes, the performance of TopRM3-L varies slightly. In addition, these results indicate that in order to obtain good overall performance, it is recommended to set the number of feedback files to 20. Figure 2 Indicates that the retrieval effect of TopRM3-L on all 5 publicly available TREC datasets is similar, and it is recommended that the feedback coefficient α be around 0.6. Brief Description of the Drawings
[0032] Figure 1 Shows the sensitivity of TopRM3-L to the number of feedback files |F|.
[0033] Figure 2 Shows the sensitivity of TopRM3-L to the feedback coefficient α. Detailed Description of the Invention
[0034] The present invention mainly uses five publicly available datasets of TREC to test the proposed model to verify the effectiveness of the proposed method. Specifically, it includes DISK1&2 with query ordinals from 51 to 150, DISK4&5 with query ordinals from 401 to 450, WT2G with query ordinals from 401 to 450, WT10G with query ordinals from 451 to 550, and GOV2 with query ordinals from 701 to 850. The scales and types of these datasets are different. In the experiments of the present invention, only the title field of TREC queries is retrieved because users of search engines always enter short enough content in their queries to represent the query intent. The present invention deletes the queries without evaluation criteria, also deletes the standard English stop words, and performs stemming on each word using the Porter English stemmer. Finally, the present invention uses the official evaluation metric of TREC, i.e., Mean Average Precision (MAP), to evaluate the effectiveness of the proposed model. All statistical tests are based on the Wilcoxon paired signed-rank test.
[0035] In the topic-based experiments of the present invention, the proposed method is compared with LM, BM25, and two typical PRF models, RM3 and Rocchio model. Although there are other PRF methods, the model of the present invention is proposed based on RM3 and Rocchio, so other PRF methods are not considered here. For fairness, the present invention uses the general settings in the IR field for the parameters of both the baseline model and the proposed model. In LM, the Dirichlet smoothing parameter is set to μ = 1000, and the paper
[23] shows that for most collections, the best MAP value is achieved. In BM25, b, k1, and k3 are set to 0.35, 1.2, and 8.0 respectively, and the paper
[24] shows that it can provide the best MAP for most collections. For the parameters in the PRF model, the number of expansion terms is fixed at 30. The parameter of the top-ranked documents traversed is |F| ∈ {10, 20, 30, 50}, the interpolation parameter α in the relevance model ∈ {0.0, 0.1,..., 1.0}, and β in the Rocchio model ∈ {0.0, 0.1,..., 1.0}. In LDA
[11] the number of topics is recommended to be 5, 10, and 20
[10] . All experimental results are evaluated through two-fold cross-validation. The TREC queries are divided according to the parity of the query coefficients. The parameters are trained on one query set and applied to evaluate the other query set, and vice versa.
[0036] In addition, to obtain good overall performance, it is recommended to set the number of feedback files to 20, and the feedback coefficient α is recommended to take a value around 0.6.
[0037] Table 1 shows all the experimental results of comparing the MAP values of the baseline model, TopRM3, and TopRoc on 5 publicly available TREC datasets. Among them, the symbol "*" indicates a statistically significant improvement (p < 0.05) compared to the corresponding PRF model at the 0.05 level according to the Wilcoxon paired signed-rank test. The percentage in parentheses is the improvement ratio compared to the corresponding PRF model. The best results in each group are marked in bold. RM3 represents the relevance model with LM as the first-pass retrieval model, and Rocchio represents BM25 + Rocchio.
[0038] Table 2 compares the proposed TopRoc with TS-COS, which is the most effective variant of TopPRF and is considered one of the most effective and state-of-the-art PRF models. For a fair comparison, the improvement percentages of them with the corresponding Rocchio model are used in the table. The comparison of the improvement percentages of TopRoc-L and TS-COS is shown in Table 2, indicating that TopRoc-L has a greater improvement.
[0039] Table 1 compares the MAP values of the baseline model, TopRM3, and TopRoc
[0040]
[0041] Table 2 compares the improvement percentages of TopRoc-L and TS-COS, indicating that TopRoc-L has a greater improvement
[0042]
Claims
1. A document topic relevance model in pseudo-relevance feedback, characterized in that Estimate the reliability of feedback documents by introducing the correlation between topic-based feedback documents into the PRF model; The PRF model is a relevance model. By introducing the relevance between feedback documents based on topics into the PRF model, a topic-based relevance model is constructed. The model is: P T (w|R) ∝ Σ D∈F P(w|D)·P T (D|F); P(w|D) is the document language model, and P T (D|F) is the topic-based document relevance. D is a document in the feedback document set F, w is the candidate term, and R represents relevance; In the theme-based relevance model where P T (z|D) is the topic distribution of the top k feedback document sets F in the first-pass retrieval in the topic space, TS(D) represents topic similarity, D i and D j are the i-th document and the j-th document respectively, and z is the topic; Or, in the topic-based relevance model where P T (z|D) is the topic distribution of the top k feedback document sets F in the first-pass retrieval in the topic space.
2. A document-topic relevance model in pseudo-relevance feedback, characterized in that, Estimate the reliability of feedback documents by introducing the correlation between topic-based feedback documents into the PRF model; the PRF model is the Rocchio model; construct a topic-based Rocchio model by introducing the correlation between topic-based feedback documents into the PRF model, and the model is specifically described as follows: (1) All documents are ranked for a given query using a specific IR model; BM25 is used in the first-pass retrieval, and the top |F| documents ranked are determined as the pseudo-relevant set F; (2) Each candidate term in the top |F| documents is assigned an expansion weight; the expansion weight is the dot product of the weights provided by the weighting model and the topic-based document correlation, and the weighting model is the TF-IDF model; (3) The vector of the query term weights is a linear combination of the initial query term weights and the extended weights, and its formula is as follows: Q1 = α·Q0 + β·∑ D∈F r(D)·P T (D|F), where Q0 and Q1 represent the original query vector and the query vector generated after one iteration respectively, α and β are adjustment parameters that control the degree of dependence on the original query vector and the feedback information, r(D) is the TF-IDF weight vector of the feedback document D, F is the feedback document set of PRF, and P T (D|F) is the topic-based document relevance; In the topic-based Rocchio model where P T (z|D) is the topic distribution in the topic space of the top k feedback document sets F in the first-pass retrieval; Or, in the topic-based Rocchio model where P T (z|D) is the topic distribution of the top k feedback document sets F in the first-pass retrieval in the topic space.
3. The document topic relevance model in a pseudo-relevance feedback according to claim 2, characterized in that α is fixed at 1.
4. The document topic relevance model in a pseudo-relevance feedback according to claim 2, wherein P T (D|F) follows a uniform distribution.