Sentence vector loop-based retrieval correlation model and construction method thereof
Through the search correlation model construction method based on sentence vector loop, the cross attention mechanism is used to process the interaction between query and document, and a task-specific embedded vector is generated, which solves the problems of existing models in interactive matching and computing resource occupation, and achieves more accurate and efficient correlation calculation.
Patent Information
- Application Number
- CN202510077157.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing search correlation model does not match the training form of the original pedestal model in the interactive matching between the query and the document, which easily destroys the basic semantic information of the original pedestal model, and consumes a lot of computing resources, which cannot carry out downstream tasks specific training, and lacks semantic calculations of the overall sentence vector, resulting in the easy fall into local optimality.
The search correlation model construction method based on sentence vector loop is adopted. By obtaining the query and document, sequence text is generated and pre-trained model is input to obtain vector representations, the cross attention mechanism is used to interact with the query and document, task-specific task embedding vectors are generated, and correlation scores are calculated.
When the query fully interacts with the document, task-specific embedding vectors are introduced, which improves the accuracy and efficiency of correlation calculations, reduces the use of computing resources, and can conduct downstream tasks specific training, avoids local optimal sorting.
Smart Images

Figure CN120011539A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic data processing, and in particular relates to a retrieval relevance model based on sentence vector loop and a construction method thereof. Background Art
[0002] In a retrieval system, re-ranking is to re-score and sort the query and the documents selected by the rough ranking. Relevance is the most important indicator to measure the semantics of the query and the document. The relevance model scores the semantic relevance between the query and different documents to determine which document corresponding to the current query better meets the user's needs.
[0003] With the rise of retrieval enhancement generation tasks, how to find documents that are more relevant to user queries in a large number of documents has become increasingly important, making re-ranking technology re-enter people's field of vision. The re-ranking task plays the role of an intelligent junk text filter in retrieval enhancement generation, filtering out irrelevant documents and inputting the most relevant documents into the large model as pre-prompt words.
[0004] The main rearrangement technologies include rule-based rearrangement, machine learning rearrangement models, linear models, tree models, end-to-end deep learning neural network models, etc. For example, based on the BERT model, a classification layer is added. The classification layer mainly includes the TextMatching model that extracts features from adjacent words based on CNN, and the Esim model that uses the mutual representation between queries and documents to improve the accuracy of results. Rearrangement technologies based on pre-training, for example, secondary pre-training of word embedding semantic similarity based on masked language model pre-training, such as UniLM, SimBERT, Sentence BERT, ERNIE-Gram, BGE, etc., can map queries and documents into sentence vectors, and then calculate text distances such as Euclidean distance, cosine similarity, and edit distance to give a similarity score between the two, and then re-rank.
[0005] However, the prior art has the following technical problems:
[0006] It mainly focuses on interactive matching between queries and documents, which is inconsistent with the training form of the original base model and easily destroys the basic semantic information obtained after pre-training of the original base model.
[0007] The correlation model based on the dual-tower structure occupies more computing resources (GPU) when reasoning;
[0008] Existing retrieval-based pre-training models can only utilize and mine the semantic knowledge obtained by pre-training the base model, and cannot perform training specific to downstream tasks (e.g., relevance tasks);
[0009] Existing relevance models mostly perform interactive operations at the word vector level, lacking semantic calculations on the overall sentence vector and between sentence vectors. In the retrieval and re-ranking stage, queries and documents are mostly sorted according to their relevance scores, which easily causes the sorting to fall into local optimality. Summary of the invention
[0010] In view of this, some embodiments disclose a method for constructing a retrieval relevance model based on a sentence vector cycle, including:
[0011] S1. Obtaining a query and a corresponding document; the document includes a title of a web page and content extracted from the web page;
[0012] S2, marking the gear position of the document;
[0013] S3, generate sequence text based on query and document;
[0014] S4, inputting the sequence text into the pre-trained model to obtain a vectorized representation of the sequence text, and obtaining a text classification vector, a query vector, a title vector, a title end vector, a content vector, and a content end vector;
[0015] S5, create a vector of the hidden layer dimension of the aligned pre-trained model output as the first vector of the retrieval task;
[0016] S6. Perform cross-attention mechanism interaction processing on the first retrieval task vector and the query vector: use the first retrieval task vector as the query of the cross-attention mechanism, and use the query vector as the key and value of the cross-attention mechanism to obtain the second retrieval task vector;
[0017] S7, perform cross-attention mechanism interaction processing on the second retrieval task vector and the document vector: concatenate the title vector and the content vector into a document vector, use the second retrieval task vector as the query of the cross-attention mechanism, and use the document vector as the key and value of the cross-attention mechanism to obtain the third retrieval task vector;
[0018] S8. Concatenate the third vector of the retrieval task with the text classification vector, multiply it by the task weight, obtain the unnormalized logarithmic probability of the calculation score, and use the Sigmoid function to map to obtain a floating-point score, which is the relevance calculation score; specifically expressed as:
[0019] Score = Sigmoid(W final Concat([cls]′;V 3 ))
[0020] Among them, W final The task weights are randomly initialized from a normal distribution, [cls] ′ is the text classification vector, V 3is the third vector of the retrieval task, Score is a floating point score between 0 and 1, and Concat is a vector concatenation operation.
[0021] Furthermore, some embodiments disclose a method for constructing a retrieval relevance model based on a sentence vector loop. In step S2, the title and content of the query and document are labeled and divided into four levels: excellent, medium, poor and irrelevant. The level identifier corresponding to the document is obtained, and the level identifier is a value between 0 and 1.
[0022] Some embodiments disclose a method for constructing a retrieval relevance model based on a sentence vector loop. In step S3, the generated sequence text is represented as: [cls] query [sep] title [sep] content [sep]; at the same time, different text type IDs are assigned to [cls], query [sep], title [sep], and content [sep] to distinguish different text types; wherein [cls] is a sequence semantic symbol, and [sep] is a sequence end symbol.
[0023] Some embodiments disclose a method for constructing a retrieval relevance model based on a sentence vector loop. In step S4, the sequence text [cls] query [sep] title [sep] content [sep] is input into a retrieval formula pre-training model to obtain a vectorized representation of the sequence text, namely, a text classification vector [cls]', a query vector query', a title vector title', a title end vector [sep1]', a content vector snippet', and a content end vector [sep2]'.
[0024] In some embodiments of the method for constructing a retrieval relevance model based on sentence vector loop, in step S6, the cross attention mechanism is a multi-head cross attention mechanism, including three parts: query sequence, key sequence, and value sequence; step S6 specifically includes:
[0025] Initialize the weights W of the normal distribution q , W k , W v , W;
[0026] Calculate Q, K, and V according to the following formula:
[0027] Q=W q ·V 1
[0028] K=W k ·query ′
[0029] V=W v ·query ′
[0030] Among them, Q is the query vector, K is the key vector, V is the value vector, and V 1 is the first vector of the retrieval task, W q is the weight of Q, W k is the weight of K, W v is the weight of V; query ′ is the query vector;
[0031] Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head respectively; the calculation method includes:
[0032] The dot product of Q and K is divided by the square root of the hidden dimension d to calculate the similarity score between the first vector of the retrieval task and each word in the query vector;
[0033] After softmax processing, the attention weight of the first vector of the retrieval task to each word in the query vector is obtained;
[0034] Multiply the attention weight of the corresponding word by the value vector V and add them together to get the final single-head vector expression Head i ; The single-head calculation process is expressed as:
[0035]
[0036] All single head vectors Head i Concatenate and multiply by the spatial weight W to get the output vector V 2 , expressed as:
[0037] V 2 =concat(Head 1 ;…;Head i )·W
[0038] Among them, V 2 As the second vector for retrieval tasks.
[0039] In some embodiments, the method for constructing a retrieval relevance model based on sentence vector loop disclosed in step S7 specifically includes:
[0040] Initialize the weights W of the normal distribution q , W k , W v , W;
[0041] Calculate Q, K, and V according to the following formula:
[0042] Q=W q ·V 2
[0043] K=W k concat(title ′;snippet ′ )
[0044] V=W v concat(title ′ ;snippet ′ )
[0045] Among them, Q is the query vector, K is the key vector, V is the value vector, and V 2 is the second vector of the retrieval task, W q is the weight of Q, W k is the weight of K, W v is the weight of V, title' is the title vector, snippet' is the content vector, and concat is the vector concatenation operation;
[0046] Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head separately, including:
[0047] The dot product of Q and K is scaled by the square root of the hidden dimension d to calculate the similarity score between the second vector of the retrieval task and each word in the document vector;
[0048] After softmax processing, the attention weight of the second vector of the retrieval task on each word in the document vector is obtained;
[0049] Multiply the attention weight of the corresponding word segment by the value vector and add them together to get the final single-head vector expression Head i ; The calculation process is expressed as:
[0050]
[0051] All Head i Concatenate and multiply by the spatial weight W to get the output vector V 3 , expressed as:
[0052] V 3 =concat(Head 1 ;…;Head i )·W
[0053] Among them, V 3 As the third vector of the retrieval task, concat is a vector concatenation operation.
[0054] The method for constructing a retrieval relevance model based on sentence vector loop disclosed in some embodiments further includes step S9, loss function calculation, specifically including: using a binary cross entropy loss function BCEloss calculated by probability, the calculation formula is:
[0055]
[0056] Among them, L i It is the file position identifier, with a value between 0 and 1. i It is the floating point score Score output by the correlation model.
[0057] In some embodiments, the method for constructing a retrieval relevance model based on sentence vector loop disclosed in step S3 includes:
[0058] Fixed query length;
[0059] The query, document title and summary are fed into the encoder-only base model to interactively obtain sequence text.
[0060] On the other hand, some embodiments disclose a retrieval relevance model based on sentence vector loop, which is obtained by the retrieval relevance model construction method based on sentence vector loop disclosed in an embodiment of the present invention.
[0061] The retrieval relevance model obtained by the retrieval relevance model construction method based on sentence vector loop disclosed in the embodiment of the present invention is a model based on a single tower structure. When the query and the document fully interact, a new task-specific task embedding vector is introduced to summarize the query, and the query is summarized into a single-word vector - the query sentence vector. Then, this query sentence vector is used to search for documents, and the correlation between the document and the query sentence vector is calculated to give a corresponding correlation score. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 A flow chart of a method for constructing a retrieval relevance model based on a sentence vector loop disclosed in some embodiments;
[0063] Figure 2 Inputting the BGE pre-trained model into the text sequence disclosed in some embodiments to obtain the relevant text vector model structure;
[0064] Figure 3 The multi-head cross-attention mechanism model structure disclosed for some embodiments. DETAILED DESCRIPTION
[0065] The word "embodiment" used herein as an "exemplary" does not necessarily mean that any embodiment described is superior or better than other embodiments. Performance index tests in the embodiments of the present invention are performed using conventional test methods in the art unless otherwise specified. It should be understood that the terms described in the embodiments of the present invention are only used to describe specific implementation methods and are not intended to limit the contents disclosed in the embodiments of the present invention.
[0066] Unless otherwise specified, the technical and scientific terms used herein have the same meanings as commonly understood by ordinary technicians in the technical field to which the embodiments of the present invention belong; other experimental methods and technical means not specifically specified in the embodiments of the present invention refer to experimental methods and technical means commonly used by ordinary technicians in the field.
[0067] The terms "substantially" and "approximately" used herein are used to describe small fluctuations. For example, they can refer to less than or equal to ±5%, such as less than or equal to ±2%, such as less than or equal to ±1%, such as less than or equal to ±0.5%, such as less than or equal to ±0.2%, such as less than or equal to ±0.1%, such as less than or equal to ±0.05%. The numerical data represented or presented in the range format herein are used only for convenience and brevity, and should therefore be flexibly interpreted as including not only the values clearly listed as the limits of the range, but also all independent values or sub-ranges contained in the range. For example, the numerical range of "1-5%" should be interpreted as including not only the clearly listed values of 1% to 5%, but also the independent values and sub-ranges within the range shown. Therefore, independent values such as 2%, 3.5% and 4% and sub-ranges such as 1%-3%, 2%-4% and 3%-5% are included in this numerical range. This principle also applies to the range of only one numerical value. In addition, such an interpretation applies regardless of the width of the range or the characteristics described.
[0068] In this document, including in the claims, transitional words such as "comprises," "includes," "with," "having," "containing," "involving," "accommodating," etc. are understood to be open-ended, i.e., meaning "including but not limited to." Only the transitional words "consisting of" and "composed of" are closed transitional words.
[0069] In order to better illustrate the present invention, numerous specific details are provided in the specific examples below. It should be understood by those skilled in the art that the present invention can be implemented without certain specific details. In the embodiments, some methods, means, instruments, equipment, etc. well known to those skilled in the art are not described in detail in order to highlight the gist of the present invention.
[0070] Under the premise of no conflict, the technical features disclosed in the embodiments of the present invention may be arbitrarily combined, and the resulting technical solutions belong to the contents disclosed in the embodiments of the present invention.
[0071] In some implementations, a method for constructing a retrieval relevance model based on a sentence vector loop includes:
[0072] S1. Obtain a query and a corresponding document; the document includes the title of a web page and the content extracted from the web page; usually, a query word query and related web pages are obtained, and the relevant content of the web page is extracted as a document content snippet based on the information of the query word query, and the title of the web page title and the document content snippet are together used as a document doc corresponding to the query word; usually, multiple related web pages are obtained based on the same query word, and multiple corresponding documents doc are obtained;
[0073] S2. Label the document with a grade mark. Usually, the query word query and the corresponding document title and content can be labeled and scored manually or based on artificial intelligence models, such as chatGPT, Qwen and other large models. Generally, the labeling and scoring can be divided into four grade marks, and the four grade marks can be further converted into labels between 0 and 1. For example, the four grades are excellent, medium, poor and irrelevant, among which excellent means a complete match, and the document fully meets the main intention of the query; medium means a moderate match, and the secondary requirements of the query are largely met by the document, or the main requirements of the query are partially met by the document; poor means a low match, the query and the document have a low degree of correspondence, and only a few contents are relevant; irrelevant means that the query and the document have no connection at all;
[0074] S3. Generate sequence text based on query and document. Usually, the generated sequence text is expressed as: [cls] query [sep] title [sep] content [sep]. At the same time, different text type IDs are assigned to [cls], query [sep], title [sep], content [sep] to distinguish different text types. Among them, [cls] is a sequence semantic symbol, and [sep] is a sequence end symbol.
[0075] In some embodiments, the length of the query is fixed; the query, the title and the abstract of the document are fed into a retrieval-based pre-trained model to interactively obtain a sequence text;
[0076] S4. Input the sequence text into the pre-training model to obtain the vectorized representation of the sequence text, and obtain the text classification vector, query vector, title vector, title end vector, content vector and content end vector; usually, a text embedding model can be used as a retrieval pre-training base model, such as a BGE model, a BERT model, etc.;
[0077] In some embodiments, the BGE model is selected as the pre-training model, such as Figure 2As shown, the sequence text [cls] query [sep] Title [sep] Snippet [sep] is input into the BGE model to obtain the vectorized representation of the sequence text, namely, the text classification vector [cls]', the query vector query', the title vector title', the title end vector [sep1]', the content vector snippet', and the content end vector [sep2]';
[0078] The BGE model is a text embedding model proposed by Beijing Zhiyuan Artificial Intelligence Research Institute, which aims to convert text into low-dimensional dense vectors for efficient calculation and analysis. The BGE model structure is consistent with RoBert, which is a bert-like model. It uses the sequence semantic symbol [cls] as the text vector and supports search-related tasks such as retrieval, re-shooting, clustering, and classification.
[0079] S5. The newly created alignment pre-trained model outputs a vector of hidden layer dimension as the first vector of the retrieval task. Usually, the newly created alignment pre-trained model outputs a vector of a token size of hidden layer dimension hidden size, which is a task-specific embedding vector specific to the downstream task, and is used as the first vector of the retrieval task corresponding to the query word query and the sequence text.
[0080] S6. Perform cross-attention interaction on the first retrieval task vector and the query vector: use the first retrieval task vector as the query of the cross-attention mechanism, and use the query vector as the key and value of the cross-attention mechanism to obtain the query-specific embedding vector as the second retrieval task vector. During the interaction process, the cross-attention mechanism calculates the similarity weights between the first retrieval task vector and each word in the query vector, so the obtained second retrieval task vector summarizes the weights of each word in the query vector.
[0081] In some embodiments, the cross-attention mechanism is a multi-head cross-attention mechanism. The multi-head cross-attention mechanism cross-attention is a natural language processing mechanism for processing dependencies between two different sequences, and includes three parts: a query vector Q, a key vector K, and a value vector V; wherein the query vector comes from one modality, and the key-value vector comes from another modality; this cross-attention mechanism allows the model to capture element-level dependencies between different modalities, and obtain relevant information by calculating the similarity between the query and the key value;
[0082] In some embodiments, Figure 3 As shown in Figure 1, the cross-attention mechanism interaction process between the first vector of the retrieval task and the query vector includes:
[0083] Initialize the weights W of the normal distribution q , W k , W v , W;
[0084] Calculate Q, K, and V according to the following formula:
[0085] Q=W q ·V 1
[0086] K=W k ·query ′
[0087] V=W v ·query ′
[0088] Among them, Q is the query vector, K is the key vector, V is the value vector, and V 1 is the first vector of the retrieval task, W q is the weight of Q, W k is the weight of K, W v is the weight of V; query ′ is the query vector;
[0089] Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head separately; including:
[0090] The dot product of Q and K is divided by the square root of the hidden dimension d to calculate the similarity score between the first vector of the retrieval task and each word in the query vector;
[0091] After softmax processing, the attention weight of the first vector of the retrieval task to each word in the query vector is obtained;
[0092] Multiply the attention weight of the corresponding word by the value vector v and add them together to get the final single-head vector expression Head i ; The single-head calculation process is expressed as:
[0093]
[0094] All single head vectors Head i Concatenate and multiply by the spatial weight W to get the output vector V 2 , expressed as:
[0095] V 2 =concat(Head 1 ;…;Head i )·W
[0096] Among them, V 2As the second vector of the retrieval task, task embedding2; concat is a vector concatenation operation;
[0097] S7, perform cross-attention mechanism interaction processing on the second retrieval task vector and the document vector: concatenate the title vector and the content vector into a document vector, use the second retrieval task vector as the query of the cross-attention mechanism, and use the document vector as the key and value of the cross-attention mechanism to obtain the third retrieval task vector;
[0098] In some embodiments, Figure 3 As shown in Figure 1, the cross-attention mechanism interaction process between the second vector of the retrieval task and the document vector includes:
[0099] Initialize the weights W of the normal distribution q , W k , W v , W; usually, initialize the weights W of the normal distribution q , W k , W v , W can use the same initialization weight as in step S6;
[0100] Calculate Q, K, and V according to the following formula:
[0101] Q=W q ·V 2
[0102] K=W k concat(title ′ ;snippet ′ )
[0103] V=W v concat(title ′ ;snippet ′ )
[0104] Among them, Q is the query vector, K is the key vector, V is the value vector, and V 2 is the second vector of the retrieval task, W q is the weight of Q, W k is the weight of K, W v is the weight of V; title' is the title vector, snippet' is the content vector; concat is the vector concatenation operation;
[0105] Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head separately, including:
[0106] The dot product of Q and K is scaled by the square root of the hidden dimension d to calculate the similarity score between the second vector of the retrieval task and each word in the document vector;
[0107] After softmax processing, the attention weight of the second vector of the retrieval task on each word in the document vector is obtained;
[0108] Multiply the attention weight of the corresponding word segment by the value vector and add them together to get the final single-head vector expression Head i ; The calculation process is expressed as:
[0109]
[0110] All Head i Concatenate and multiply by the spatial weight W to get the output vector V 3 , expressed as:
[0111] V 3 =concat(Head 1 ;…;Head i )·W
[0112] Among them, V 3 As the third vector of the retrieval task, task embedding3, concat is a vector concatenation operation;
[0113] S8. Concatenate the third vector of the retrieval task with the text classification vector, multiply it by the task weight, obtain the unnormalized logarithmic probability of the calculation score, and use the Sigmoid function to map to obtain a floating-point score, which is the relevance calculation score; specifically expressed as:
[0114] Score = Sigmoid(W final Concat([cls]′;V 3 ))
[0115] Among them, W final The task weights are randomly initialized from a normal distribution, [cls] ′ is the text classification vector, V 3 is the third vector of the retrieval task, Score is a floating point score with a value between 0 and 1;
[0116] S9, loss function calculation, specifically including: using probability calculation of the binary cross entropy loss function BCEloss, the calculation formula is:
[0117]
[0118] Among them, L i It is the file position identifier, with a value between 0 and 1. i It is the floating point score Score output by the correlation model.
[0119] Some embodiments disclose a retrieval relevance model based on sentence vector loop, which is obtained by a retrieval relevance model construction method based on sentence vector loop disclosed in an embodiment of the present invention.
[0120] The method for constructing a retrieval relevance model based on sentence vector cycle disclosed in the embodiment of the present invention is a single-tower structure model. When the query and the document doc fully interact through the pre-trained model, a new task-specific task embedding task-specific embedding is introduced to summarize the query query, and the query query is summarized into a vector of a single word token - the query sentence vector query-specific embedding, and then the query sentence vector is used to search for the document doc, and the correlation between the document doc and the query sentence vector query-specific embedding is calculated to give a corresponding correlation score; the overall design refers to the sequence input method of the recurrent neural network RNN, and the overall sentence vector is used as the input of each time step, and the sentence vector calculation method under each cycle time step is the cross attention mechanism, which can calculate the weight of each word vector token in the sentence with the correlation task, and then obtain the final vector expression;
[0121] The query length is fixed, and the title and summary of the document doc are concatenated into the encoder-only retrieval pre-training model. The query and the document doc are simultaneously passed through the self-attention encoding module of each layer of the basic model, and can express each other. Under the same computing resources (such as the graphics processing unit GPU), the single-tower model can better save computing resources. At the same time, the query and the document are interacted in advance at the input end, and the self-attention weight expression calculation is performed in advance between the two, which can maximize the use of the pre-trained semantic space of the retrieval pre-training model.
[0122] The classification layer draws on the sequence modeling method of the recurrent neural network (RNN), replacing the elements with sentence vectors that can better express global information rather than single word tokens. This method can generate subsequent elements based on the previous elements, abandoning the interactive matrix between the query vector and the document vector, reducing the space occupied by the GPU memory for storage and calculation, and is more intelligent in capturing local information. It is also conducive to the embedding of more NLP technologies, such as term weight information.
[0123] A new randomly initialized, downstream task-specific embedding is introduced as the first vector of the retrieval task, which is specifically used to match queries and documents. Since the pre-trained model does not match the task form of the downstream task, the interference of the pre-trained model on the downstream task can be reduced. At the same time, the word embedding information given by the pre-training task can be used to calculate the relevant score of the relevance model.
[0124] The cross-attention mechanism is used to calculate the intermediate hidden features hiddenstates of the classification layer. Compared with the convolutional neural network CNN, this method pays more attention to global information, and the model can obtain the weight of each word token under the query for the relevance task itself through data training; the use of the cross-attention mechanism can enhance the expressiveness of the model by calculating the different weights of each word in the query, capture the complex relationship between the input data, and make the model more generalized. It can extract relevant information from different inputs through learning to help the model better generalize to unseen data; and the cross-attention mechanism has better interpretability, which can intuitively show how the model pays attention to the importance weights of each word in the query and document, which is convenient for better evaluation of the learning degree of the model when reasoning the model;
[0125] Two cross-attention interaction calculations are performed at the classification layer. The first interaction calculation is performed using the first retrieval task vector and the query vector. The vector result obtained by the first attention mechanism calculation is the query-based task embedding as the second retrieval task vector, which is input into the second cross-attention mechanism calculation together with the document to obtain the third retrieval task vector. This belongs to the cross-attention mechanism calculation of different information sources under natural language processing.
[0126] In terms of optimization, the BCE loss used is a probability calculation loss, which converts the correlation between the query and the document into a probability value between 0 and 1. It can directly optimize the probability value of the model output, so that the output learned by the model is close to the probability representation in the true probability distribution. It is more sensitive to small changes in probability without losing computational stability or imposing excessive penalties on extreme values, which helps the model to more accurately predict the probability of the current sample.
[0127] The technical solutions disclosed in the embodiments of the present invention and the technical details disclosed in the embodiments are merely illustrative of the inventive concept of the present invention and do not constitute a limitation on the technical solutions of the embodiments of the present invention. Any conventional changes, replacements or combinations of the technical details disclosed in the embodiments of the present invention have the same inventive concept as the present invention and are within the protection scope of the claims of the present invention.
Claims
1. A retrieval relevance model construction method based on sentence vector loop, characterized in that: include: S1. Obtaining a query and a corresponding document; the document includes a title of a web page and content extracted from the web page; S2, marking the gear position of the document; S3, generate sequence text based on query and document; S4, inputting the sequence text into the pre-trained model to obtain a vectorized representation of the sequence text, and obtaining a text classification vector, a query vector, a title vector, a title end vector, a content vector, and a content end vector; S5, create a vector of the hidden layer dimension of the aligned pre-trained model output as the first vector of the retrieval task; S6. Perform cross-attention mechanism interaction processing on the first retrieval task vector and the query vector: use the first retrieval task vector as the query of the cross-attention mechanism, and use the query vector as the key and value of the cross-attention mechanism to obtain the second retrieval task vector; S7, perform cross-attention mechanism interaction processing on the second retrieval task vector and the document vector: concatenate the title vector and the content vector into a document vector, use the second retrieval task vector as the query of the cross-attention mechanism, and use the document vector as the key and value of the cross-attention mechanism to obtain the third retrieval task vector; S8. Concatenate the third vector of the retrieval task with the text classification vector, multiply it by the task weight, obtain the unnormalized logarithmic probability of the calculation score, and use the Sigmoid function to map to obtain a floating-point score, which is the relevance calculation score; specifically expressed as: Score=Sigmoid(W final ·Concat([cls]′;V3)); Among them, W final is the task weighted weight randomly initialized from the normal distribution, [cls]′ is the text classification vector, V3 is the third vector of the retrieval task, Score is a floating point score with a value between 0 and 1, and Concat is a vector concatenation operation.
2. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 1, characterized in that: In step S2, the query and the title and content of the document are marked and divided into four levels: excellent, medium, poor and irrelevant, and the level identification corresponding to the document is obtained, and the level identification is a value between 0 and 1.
3. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 2, characterized in that: In step S3, the generated sequence text is represented as: [cls] query [sep] title [sep] content [sep]; at the same time, different text type IDs are assigned to [cls], query [sep], title [sep], content [sep] to distinguish different text types; Among them, [cls] is the sequence semantic symbol, and [sep] is the sequence end symbol.
4. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 3 is characterized in that: In step S4, the sequence text [cls] query [sep] title [sep] content [sep] is input into the retrieval pre-training model to obtain the vectorized representation of the sequence text, namely, the text classification vector [cls]', the query vector query', the title vector title', the title end vector [sep1]', the content vector snippet', and the content end vector [sep2]'.
5. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 4, characterized in that: In step S6, the cross attention mechanism is a multi-head cross attention mechanism, which includes three parts: query sequence, key sequence, and value sequence; step S6 specifically includes: Initialize the weights W of the normal distribution q , W k , W v , W; Calculate Q, K, and V according to the following formula: Q=W q ·V1 K=W k ·query′ V=W v ·query′ Among them, Q is the query vector, K is the key vector, V is the value vector, V1 is the first vector of the retrieval task, and W q is the weight of Q, W k is the weight of K, W v is the weight of V; query′ is the query vector; Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head respectively; the calculation method includes: The dot product of Q and K is divided by the square root of the hidden dimension d to calculate the similarity score between the first vector of the retrieval task and each word in the query vector; After softmax processing, the attention weight of the first vector of the retrieval task to each word in the query vector is obtained; Multiply the attention weight of the corresponding word by the value vector V and add them together to get the final single-head vector expression Head i ; The single-head calculation process is expressed as: All single head vectors Head i Concatenate and multiply by the spatial weight W to get the output vector V2, expressed as: V2=concat(Head1;…;Head i )·W Among them, V2 is the second vector of the retrieval task.
6. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 5, characterized in that: Step S7 specifically includes: Initialize the weights W of the normal distribution q , W k , W v , W; Calculate Q, K, and V according to the following formula: <h2 style=";text-align:left;direction:ltr">Q=W<h2 style=";text-align:left;direction:ltr"> q <h2 style=";text-align:left;direction:ltr"> V2 K=W k ·concat(title′;snippet′) V=W v ·concat(title′;snippet′) Among them, Q is the query vector, K is the key vector, V is the value vector, V2 is the second vector of the retrieval task, and W q is the weight of Q, W k is the weight of K, W v is the weight of V; title' is the title vector, snippet' is the content vector; Concat is the vector concatenation operation; Divide Q, K, and V into multiple heads, here n heads, and calculate the attention of each head separately, including: The dot product of Q and K is scaled by the square root of the hidden dimension d to calculate the similarity score between the second vector of the retrieval task and each word in the document vector; After softmax processing, the attention weight of the second vector of the retrieval task on each word in the document vector is obtained; Multiply the attention weight of the corresponding word segment by the value vector and add them together to get the final single-head vector expression Head i ; The calculation process is expressed as: All Head i Concatenate and multiply by the spatial weight W to get the output vector V3, expressed as: V3=concat(Head1;…;Head i )·W Among them, V3 is the third vector of the retrieval task, and Concat is the vector concatenation operation.
7. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 6, characterized in that: The method further includes step S9, loss function calculation, which specifically includes: using a binary cross entropy loss function BCEloss calculated by probability, and the calculation formula is: Among them, L i It is the file position identifier, with a value between 0 and 1. i It is the floating point score Score output by the correlation model.
8. The method for constructing a retrieval relevance model based on sentence vector loop according to claim 1, characterized in that: Step S3 includes: Fixed query length; The query, document title and summary are fed into the encoder-only base model to interactively obtain sequence text.
9. The retrieval relevance model based on sentence vector loop is characterized by: Obtained by the method for constructing a retrieval relevance model based on sentence vector loop as described in any one of claims 1 to 8.
Citation Information
Patent Citations
DETR-based human-object interaction detection method for human pairwise decoding interaction
CN115147931A
Object classification method and device
CN115410000A
Cross-attention system and method for fast video-text retrieval task with image clip
WO2022261570A1