A long text retrieval method and system based on Gaussian kernel function
By using Gaussian kernel functions and pre-trained language model methods in long text retrieval, the semantic similarity between paragraphs and user-retrieval contents and generating pseudo-labels is solved, and the problem of lack of paragraph-level annotation data in long text retrieval is improved, and the accuracy of the retrieval model is improved.
Patent Information
- Application Number
- CN202111512377.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-08
AI Technical Summary
In long text retrieval tasks, prior art is difficult to find paragraph-level correlation tags without additional labeling of data, and it is difficult to find a bridge between semantic similarity and user click correlation.
Using a Gaussian kernel function method, the semantic similarity between the paragraphs of long text and the contents of the user searched as pseudo-labels is used to calculate the semantic similarity between the paragraphs of long text and the contents of the user searched as pseudo-labels, and the pseudo-labels are mapped into multi-dimensional scores through different Gaussian kernel functions. Finally, the correlation score of the user searched content for the whole of long text is output through linear layer aggregation of paragraph scores.
It effectively alleviates the problem of lack of paragraph-level labeling data, enhances the correlation between semantic similarity and user click correlation, and improves the accuracy of long text retrieval models.
Smart Images

Figure CN114328863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a long text retrieval method and system, and in particular to a long text retrieval method and system based on a Gaussian kernel function, and belongs to the technical field of information retrieval. Background Art
[0002] Long text retrieval is a basic task in the field of information retrieval. Its characteristics are: the average length of the documents to be retrieved is long, and a single document may contain multiple topics. Traditional retrieval models find it difficult to locate topics related to user click intent in long texts.
[0003] In recent years, pre-trained language models have performed outstandingly in the field of information retrieval. Their powerful contextual semantic modeling capabilities enable retrieval models to better calculate the semantic similarity between user retrieval content and candidate documents, thereby improving the accuracy of the model in determining whether the two are related. However, in long text retrieval tasks, due to the limitation of input length, pre-trained language models cannot calculate the semantic similarity between user retrieval content and the entire long text.
[0004] At present, the existing technology mainly chooses to segment long texts and concatenates them with user search content in paragraph units as the input of the retrieval model. However, in the existing public datasets, the model training stage still lacks the relevance labels of paragraphs and user searches. At the same time, since semantic similarity and user click relevance are not completely equivalent, users may also click on candidate documents with lower similarity.
[0005] In summary, how to find paragraph-level relevance tags without the need for additional annotated data, and how to find a bridge between semantic similarity and user click relevance, have become urgent technical challenges facing long text retrieval. Summary of the invention
[0006] The purpose of the present invention is to solve the technical problems faced by long text retrieval, such as how to find paragraph-level relevance tags without the need for additional annotated data, and how to find a bridge between semantic similarity and user click relevance, and creatively propose a long text retrieval method and system based on Gaussian kernel function.
[0007] The innovation of this method is that it uses the semantic modeling ability of the pre-trained language model to calculate the semantic similarity between each paragraph of the long text and the user's search content as a pseudo-label of the user's click relevance. Through different Gaussian kernel functions, the pseudo-label is mapped to relevance scores of different dimensions. The linear layer is used to aggregate the scores of each paragraph of the long text to output the relevance score of the user's search content to the long text as a whole.
[0008] Beneficial Effects
[0009] Compared with the prior art, the present invention has the following advantages:
[0010] 1. The present invention uses the semantic modeling capability of the pre-trained language model to calculate the semantic similarity between the paragraph and the user's search content as a pseudo-label, which effectively alleviates the problem of lack of paragraph-level annotation data.
[0011] 2. The present invention uses a Gaussian kernel function to map pseudo-label scalars into multidimensional vectors, which can allow paragraphs with different semantic similarity levels to contribute to the relevance of user clicks, enhance the correlation between semantic similarity and user click relevance, and improve the accuracy of the long text retrieval model. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Flow chart of the method of the present invention.
[0013] Figure 2 It is a schematic diagram of the structural composition of the system of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0015] A long text retrieval method based on Gaussian kernel function, such as Figure 1 As shown, the following steps are included:
[0016] Step 1: Segment the long text.
[0017] Specifically, a length N is specified as the maximum length of each paragraph after the long text is segmented. Within the length N range, punctuation marks are used as the priority segmentation cutoff points to ensure the semantic integrity of the segmented text.
[0018] Step 2: Use the pre-trained language model to score user queries and candidate paragraphs.
[0019] Specifically, the user search content is concatenated with the candidate paragraphs, and then pre-trained, and the output sentence vector [CLS] is used as the text feature interaction vector. Then, the multi-layer perceptron MLP is used to determine the semantic similarity between the user query and the candidate paragraph as a pseudo label.
[0020] Step 3: Use the Gaussian kernel function to map the pseudo labels.
[0021] Specifically, the pseudo-label scalar is mapped to a multi-dimensional score vector through a pre-designed Gaussian kernel with different means and the same variance. Then, the vectors corresponding to different paragraphs are concatenated together to form a score matrix.
[0022] Step 4: Use the linear layer to determine the relevance of user clicks.
[0023] Specifically, the score matrix is passed through the pooling layer and then into the linear layer, and MLP is used to determine the contribution of each paragraph of the long text to the click relevance of the end user at different levels.
[0024] Furthermore, the present invention proposes a long text retrieval system based on Gaussian kernel function, such as Figure 2 As shown, it includes a pseudo-label calculation module, a Gaussian kernel mapping module and an output module.
[0025] Among them, the pseudo-label calculation module is responsible for segmenting long documents, and concatenating the obtained text paragraphs with the user search content and inputting them into the pre-trained language model to obtain the text feature interaction vector. At the same time, the text feature interaction vector is used as the input of the linear layer, and the correlation between the output user search content and each paragraph of the long text is used as the pseudo-label.
[0026] The Gaussian kernel mapping module is responsible for mapping pseudo labels from scalars to score vectors through different Gaussian kernel functions.
[0027] The output module is used to concatenate the score vectors of different paragraphs belonging to the same long text into a score matrix, average-pool the score matrix and put it into the linear layer to determine and integrate the correlation between the user's search content and the long text under different Gaussian kernel functions.
[0028] The connection relationship between the above modules is:
[0029] The output end of the pseudo label calculation module is connected to the input end of the Gaussian kernel mapping module. The output end of the Gaussian kernel mapping module is connected to the input end of the output module.
[0030] The system works as follows:
[0031] First, the long text is segmented in the pseudo-label calculation module. To ensure the completeness of the segmented paragraphs, the segmentation cutoff points are first graded according to priority, where punctuation marks have a higher priority than the specified maximum paragraph length. Then, the segmented paragraphs are concatenated with the user search content and input into the pre-trained language model to obtain the text feature interaction vector. Finally, the text feature interaction vector is put into the linear layer, and the correlation between the user search content and each paragraph of the long text is output as a pseudo-label.
[0032] In the pseudo-label calculation module, the pre-trained language model can use the BERT model to obtain the text feature interaction vector V i , as shown in formula 1:
[0033] V i =BERT(q,p j ) (1)
[0034] The value range of i is 1, 2, 3, ..., n, where n refers to the maximum number of paragraphs that a long text can be divided into; q is the user's search content, and p is the maximum value of the paragraphs that can be divided into paragraphs. j is the jth paragraph of the long text.
[0035] The linear layer refers to a fully connected neural network that maps the text feature interaction vector to correlation, as shown in Formula 2:
[0036] R=W*V i +b (2)
[0037] Among them, R represents the relevance score of the model output, W and b are model parameters, which can be solved by back propagation during model training; V i Represents the text feature interaction vector between the i-th paragraph and the user's search content.
[0038] The pre-trained language model BERT may adopt the Bidirectional Encoder Representations from Transformer proposed by Google AI in October 2018.
[0039] In the Gaussian kernel mapping module, the means and variances of different Gaussian kernels are first initialized, where the means of each Gaussian kernel are different but the variances are systematic. Then, the pseudo-labels output by the pseudo-label calculation module are put into different Gaussian kernels for mapping, and the obtained results are cascaded together to form a score vector. The Gaussian kernel function mapping is shown in Formula 3:
[0040] K(R i )=exp(-(R i -μ k ) / 2σ k 2 ) (3)
[0041] Among them, K(R i ) indicates that R i Retrieve content q and pseudo-label of the i-th paragraph for the user, μ k , σ k They represent the mean and variance of the kth Gaussian kernel respectively, and exp is the exponential function.
[0042] In the output module, the score vectors corresponding to different paragraphs of the long text are first concatenated together to obtain a score matrix. After the score matrix is averaged and pooled, it is input into the linear layer to output the final user search content and long text relevance score. Finally, MLP is used to determine the contribution of each paragraph of the long text to the final user click relevance at different levels.
Claims
1. A long text retrieval system based on Gaussian kernel function, characterized in that: It includes a pseudo-label calculation module, a Gaussian kernel mapping module and an output module; Among them, the pseudo-label calculation module is responsible for segmenting long documents, and concatenating the obtained text paragraphs with the user search content and inputting them into the pre-trained language model to obtain the text feature interaction vector; at the same time, the text feature interaction vector is used as the input of the linear layer, and the correlation between the output user search content and each paragraph of the long text is used as the pseudo-label; The Gaussian kernel mapping module is responsible for mapping pseudo labels from scalars to score vectors through different Gaussian kernel functions; The output module is used to concatenate the score vectors of different paragraphs belonging to the same long text into a score matrix, average the score matrix and put it into the linear layer to determine and integrate the relevance of the user's search content with the long text under different Gaussian kernel functions; The connection relationship between the above modules is: The output end of the pseudo-label calculation module is connected to the input end of the Gaussian kernel mapping module; the output end of the Gaussian kernel mapping module is connected to the input end of the output module; First, the long text is segmented in the pseudo-label calculation module; the segmentation cutoff points are first graded according to priority, where punctuation marks have a higher priority than the specified maximum paragraph length; then, the segmented paragraphs are concatenated with the user search content and input into the pre-trained language model to obtain a text feature interaction vector; finally, the text feature interaction vector is put into a linear layer, and the correlation between the user search content and each paragraph of the long text is output as a pseudo-label; In the pseudo-label calculation module, the pre-trained language model obtains the text feature interaction vector V i , as shown in formula 1: V i =BERT(q,p j ) (1) The value range of i is 1, 2, 3, ..., n, where n refers to the maximum number of paragraphs that a long text can be divided into; q is the user's search content, and p is the maximum value of the paragraphs that can be divided into paragraphs. j is the jth paragraph of the long text; The linear layer refers to a fully connected neural network that maps the text feature interaction vector to correlation, as shown in Formula 2: R=W*V i +b (2) Among them, R represents the relevance score of the model output, W and b are model parameters, which can be solved by back propagation during model training; V i Represents the text feature interaction vector between the i-th paragraph and the user's search content; In the Gaussian kernel mapping module, the means and variances of different Gaussian kernels are first initialized, wherein the means of each Gaussian kernel are different but the variances are systematic; then, the pseudo-labels output by the pseudo-label calculation module are put into different Gaussian kernels for mapping, and the obtained results are cascaded together to form a score vector; the Gaussian kernel function mapping is shown in Formula 3: K(R i )=exp(-(R i -m k ) / 2σ k 2 ) (3) Among them, K(R i ) indicates that R i Retrieve content q and pseudo-label of the i-th paragraph for the user, μ k , σ k They represent the mean and variance of the kth Gaussian kernel respectively, and exp is the exponential function; In the output module, the score vectors corresponding to different paragraphs of the long text are first concatenated together to obtain a score matrix; after average pooling, the score matrix is input into the linear layer to output the final user search content and long text relevance score; finally, MLP is used to determine the contribution of each paragraph of the long text to the final user click relevance at different levels.
2. A long text retrieval system based on Gaussian kernel function as claimed in claim 1, characterized in that: The pre-trained language model is the BERT model.
Citation Information
Patent Citations
Cross-border ethnic culture text retrieval method integrated with document word weight
CN112948537A