Text data enhancement method for improving vector retrieval performance

By using large language models to augment text data and filtering and integrating data based on semantic features, the problem of low data quality generated by existing data augmentation methods is solved, and the performance of vector retrieval model is improved.

CN119961436AActive Publication Date: 2025-05-09DALIAN UNIV OF TECH +1

Patent Information

Application Number
CN202510139067.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-09
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The enhanced data generated by the existing data augmentation methods in the field of vector search are not of high quality and are easily recognized by the model, resulting in poor data augmentation effect.

Method used

Data augmentation is adopted for large language models, text data is flexible with zero sample prompts, and generated data is filtered based on semantic features to ensure the quality of data augmentation. Then, the semantic information of the enhanced data is integrated using high-dimensional vector orthogonality to build guide vectors and guide the fine-tuning process of the model.

Benefits of technology

The performance of model vector retrieval is improved, the quality of the generated training data is higher, the feature representation learned by the model is more accurate, and the performance of downstream tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961436A_ABST
    Figure CN119961436A_ABST
Patent Text Reader

Abstract

The invention provides a text data enhancement method for improving vector retrieval performance, and belongs to the field of computer data analysis. The method comprises the following steps: firstly, compressing long text data by using a prompt template of a large language model, and decomposing the long text data into a plurality of short texts; in the training process, the short texts replace the original long texts to be used as training data, so that the video memory size occupied by a single piece of information is saved; in order to cope with the problem that the representation capability is reduced possibly due to the fact that the text length is shortened, a guide vector is constructed by combining a plurality of short texts from the same long text, and the guide vector is used as auxiliary information to guide the encoding process of a single short text. In this way, the adverse effect of text shortening on the model representation ability can be effectively reduced, and therefore the training effect and generalization ability of the model are improved on the premise that shorter single information is used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer data analysis, and in particular to a method for constructing a semantic space based on deep learning and enhancing text data based on a large language model. Background Art

[0002] Retrieval tasks usually refer to obtaining the set of documents that are closest to the query in a large-scale document data. At present, article retrieval technology based on dense semantic vectors has demonstrated its wide practicality in many key fields, including but not limited to open domain question answering, conversational systems, and web search.

[0003] Data augmentation is one of the classic techniques for improving the quality of information retrieval. It creates additional training data by slightly perturbing the training samples, such as rotating, translating, and scaling. Data augmentation includes a series of techniques for generating new training samples from the original samples through jitter and interference, but the class labels do not change. The purpose of data augmentation is to increase the generalization ability of the model. If the network we want to train constantly sees new, slightly changed input sample points, it can learn more robust features. However, the current data augmentation methods in the field of vector retrieval are limited to simple word and sentence transformations in the source text, such as repetition, deletion, and insertion of meaningless characters. The enhanced data generated by these methods is of low quality and can be easily recognized by the model, resulting in poor data augmentation effects.

[0004] Contrastive learning is an unsupervised or self-supervised learning method whose main goal is to learn a representation learning model by automatically constructing similar and dissimilar instances so that similar instances are closer in the projected space, while dissimilar instances are farther apart in the projected space. This method is widely used in the field of natural language processing. However, how to effectively construct similar and dissimilar instances, and how to design the model structure to follow this guiding principle, is an important issue in contrastive learning. Inappropriate instance construction may cause the model to learn incorrect feature representations, thereby affecting its performance in downstream tasks.

[0005] Based on this, the present invention proposes a method for data enhancement using a large language model to improve retrieval performance. By using zero-sample prompts to perform more flexible data enhancement on text data, and screening the data based on the semantic features when generating the data, the quality of data enhancement is ensured. After obtaining high-quality data, the present invention uses the orthogonality of high-dimensional vectors to integrate the semantic information of the enhanced data, and uses the integrated information as a guide vector to guide the fine-tuning process of the model. Ultimately, the performance of model vector retrieval is improved. Summary of the invention

[0006] The present invention aims at the problem of data enhancement in information retrieval, and proposes a method for enhancing text data using a large language model to improve retrieval performance. The method uses the prompt template of the large language model to semantically decouple the original data, and then decomposes it into several short texts; and the data generated by the large language model is filtered through semantic features to obtain enhanced data. In the training process, contrastive learning is first used for the first step of training, and then the model obtained in the first step of training is used to aggregate the enhanced data to obtain a guidance vector, and then the guidance vector is used as auxiliary information to guide the text representation process in contrastive learning to improve the performance of model vector retrieval.

[0007] The technical solution of the present invention is:

[0008] A text data enhancement method for improving vector retrieval performance includes the following steps:

[0009] S1, obtaining a data set and preprocessing the data set;

[0010] S2, using the preprocessed data set to perform comparative learning training on the pre-trained model to obtain a trained primary retrieval model;

[0011] S3, semantically decoupling the positive example documents in the preprocessed data set based on the large language model, decomposing the positive example documents into several short texts; and filtering the short text data to obtain an enhanced data set;

[0012] S4. Aggregate the data in the enhanced data set based on the primary retrieval model to construct a guidance vector, and use the guidance vector to guide the training of the primary retrieval model to obtain a final retrieval model.

[0013] Furthermore, in S1, the data set includes:

[0014] A query file, which identifies all query contents;

[0015] A file that lists the correspondence between queries and positive examples. This file can determine the correct document relationship between the query and the corresponding document.

[0016] A document collection file that identifies the contents of all documents.

[0017] Furthermore, in S1, the data set is preprocessed, specifically:

[0018] If the data set contains a correspondence file between queries and negative examples, the positive and negative documents are determined based on the correspondence files between the queries and the positive and negative examples, and the data set format is processed into (query, positive, negative); if the data set lacks a correspondence file between queries and negative examples, the retrieval model is used to retrieve similar documents in the entire document collection as query negative examples, and the positive documents are determined based on the correspondence file between the queries and the positive examples, and the data set format is processed into (query, positive, negative).

[0019] Furthermore, in S2, the training process of the primary retrieval model is as follows:

[0020] S2.1, use the pre-trained model to initialize its pre-trained checkpoint, the pre-trained model can be used as an encoder to map all queries and positive examples in the data set obtained in S1 to the same vector space;

[0021] S2.2, divide the data set obtained in S1 into several training batches according to the storage space limit of the computing device, and the training task is also divided into several batches; in any batch of training tasks, the query, positive example documents and negative example documents in the batch are first encoded into vectors at the same time;

[0022] S2.3. In the vector space of the same training batch, for any query vector, its corresponding negative examples include all documents except its own positive example document. The distances between the query vector and each positive example document vector and each negative example document vector are calculated, and the calculated distances are smoothed by the softmax function to make them distributed between (0,1). Finally, the cross entropy is used as the loss function for training to obtain a trained primary retrieval model, thereby shortening the distance between the query and the positive example and increasing the distance between the query and the negative example. The loss function formula is as follows:

[0023]

[0024] in, is the result of the distance between the query vector and the positive document vector or the negative document vector after being processed by the softmax function; y is the actual label, when the query vector and the positive document vector calculate the distance y = 1, when the query vector and the negative document vector calculate the distance, y = 0.

[0025] Furthermore, in S2.1, the pre-trained models include RetroMAE, GTE, etc.

[0026] Furthermore, the specific process of S3 is as follows:

[0027] S3.1. For a training set of the data set in S1, construct each positive example document into an instruction using a prompt template and input it into the large language model to obtain the semantically decoupled short text output of each positive example document, and then change the training set format to (query, positive example, negative example, short text set);

[0028] S3.2. For the short texts in the generated (query, positive example, negative example, short text set), the quality control index M is used to quantitatively evaluate the quality of the short texts. The quality control index M comprehensively considers the similarity of two dimensions: one is the similarity between the short text and its corresponding long text, which reflects the ability of the short text to retain the original information; the other is the similarity difference between different short texts corresponding to the same long text, which measures the degree of difference in the short text when expressing the same long text content in a diversified way; its calculation formula is as follows:

[0029] M i =∑ Aj∈Z sim(A,Ai)-∑ Ai,Aj∈Z sim(Ai,Aj)

[0030] Where Z is the short text set corresponding to the positive example A, Ai and Aj are two different short texts in the short text set Z; sim(A,Ai) represents the similarity between the positive example A and the generated short text Ai, which is the cosine similarity of the two vectors calculated after encoding the two texts A and Ai into the vector space using the primary retrieval model obtained by S2; M i represents the quality control indicator of the short text Ai;

[0031] S3.3. Set the hyperparameters K and R. For all short text sets, the quality control index of each short text data must be greater than R. At most K short texts with the largest quality control index are retained to form an enhanced data set in the format of (query, positive example, negative example, enhanced short text set).

[0032] Furthermore, in S3.1, the large language model includes GPT-4, Tongyi, etc.

[0033] Furthermore, the specific process of S4 is as follows:

[0034] S4.1. First, load the primary retrieval model obtained in S2, and divide the enhanced data set obtained in S3 into several training batches according to the storage space size limit of the computing device. The training task is therefore divided into several batches. In a batch of training tasks, the queries, positive examples, negative examples, and enhanced short text sets in the batch are all encoded into high-dimensional vectors, where the enhanced short text set does not pass through the pooling layer. Add all the short text high-dimensional vectors from the same query to obtain a concatenated vector. Then pass the concatenated vector through a fully connected layer and a residual connection layer, and project it into a high-level semantic space. Split the concatenated vector of each query according to the vector position, and use the knn algorithm to cluster and generate several phrase vectors based on the Euclidean distance between the vectors at each position. For each query vector, calculate the KL divergence between the query vector and each phrase vector, and use the phrase vector corresponding to the minimum KL divergence as the guide vector. The calculation formula of the guide vector is:

[0035] P=min u∈U (D KL (Q||u))

[0036] Among them, P represents the guidance vector, Q represents the high-dimensional vector encoded by the query, U represents the set of phrase vectors, u is a single phrase vector, and D KL represents the divergence function;

[0037] S4.2. Use the loss function L to perform training and learning to obtain the final retrieval model; the loss function L = L1 + L2, respectively as follows:

[0038] For the query and short text set, mark any short text corresponding to the query as a positive example and the rest as negative examples to calculate the cross entropy as the loss function L1;

[0039] L1=∑-ylogsim(Q,S)-(1-y)log(1-sim(Q,S))

[0040] Among them, sim represents the cosine distance between vectors; Q represents the high-dimensional vector encoded by the query; S represents the high-dimensional vector encoded by the short text; y is the actual label, when the short text is a positive example, y = 1, when the short text is a negative example, y = 0;

[0041] For the guidance vector and the positive example, the cross entropy between the guidance vector corresponding to any query and the positive example corresponding to the query and the negative example corresponding to the query is calculated as the loss function L2;

[0042] L2=∑-ylogsim(P,D)-(1-y)log(1-sim(P,D))

[0043] Among them, sim represents the cosine distance of the vector; P represents the guidance vector; D represents the vector encoded by the positive example or the negative example; y is the actual label, when D is the vector encoded by the positive example, y=1, when D is the vector encoded by the negative example, y=0.

[0044] The beneficial effects of the present invention compared with the prior art are as follows:

[0045] (1) To address the problem of not following instructions when using large language models for data compression, a quality assessment method that focuses on the consistency between the generated text and the original text and the differences between the generated texts was used to improve the quality of training data;

[0046] (2) To address the information loss problem caused by text compression, a method is adopted to construct a guidance vector by interactively generating multiple short texts to restore text information within the model, thereby reducing the loss of compressed information and improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is the overall framework diagram of the present invention;

[0048] Figure 2 is an example of an instruction in the present invention;

[0049] Figure 3 This is a flow chart of obtaining phrase vectors in the present invention. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0051] like Figure 1 As shown, the embodiment of the present invention provides a text data enhancement method for improving vector retrieval performance, the steps are as follows:

[0052] S1. Obtaining a data set and preprocessing the data set

[0053] The dataset Multi-CPR was selected and preprocessed, specifically: each document in the dataset was vector-embedded using the BGE model to obtain a vector representation of each document; then the generated vector embeddings were stored in the FAISS library developed by FacebookAI to build an efficient vector index database. In the vector index database built using FAISS, the entire document collection was retrieved for document embedding based on the approximate nearest neighbor search algorithm (ANNS) to obtain negative documents. Similar documents were retrieved from the entire document collection as query negative examples, and positive documents were determined based on the correspondence between the query and the positive example, and the dataset format was processed into (query, positive example, negative example).

[0054] S2. Use the data set obtained in S1 to perform comparative learning training on the pre-trained model to obtain a trained primary retrieval model. The specific process is as follows:

[0055] S2.1, using the pre-trained checkpoint provided by the pre-trained model RetroMAE as initialization, the pre-trained model RetroMAE can be used as an encoder to map all queries and positive examples in the data set obtained in S1 to the same vector space;

[0056] S2.2, the training batch size is set to 32; in any batch of training tasks, the query, positive document and negative document in the batch are first encoded into vectors at the same time, and the negative document corresponding to any query includes all documents except its own positive document;

[0057] S2.3. In the vector space of the same training batch, for any query vector, its corresponding negative examples include all documents except its own positive example document. The distances between the query vector and each positive example document vector and each negative example document vector are calculated, and the calculated distances are smoothed by the softmax function to make them distributed between (0,1). Finally, the cross entropy is used as the loss function for training to obtain a trained primary retrieval model, thereby shortening the distance between the query and the positive example and increasing the distance between the query and the negative example. The loss function formula is as follows:

[0058]

[0059] in, is the result of the distance between the query vector and each positive document vector or negative document vector after being processed by the softmax function; y is the actual label, when the query vector and the positive document vector calculate the distance y = 1, when the query vector and the negative document vector calculate the distance, y = 0.

[0060] S3: Based on the large language model, semantically decouple the positive example documents in the data set of S1, decompose the positive example documents into several short texts, and filter the short text data to obtain an enhanced data set. The specific process is as follows:

[0061] S3.1. For a training set of the data set in S1, each positive example document is constructed into the instruction shown in the figure using a prompt template and input into the large language model GPT-4 to obtain the semantically decoupled short text output of each positive example document, and then the training set format is changed to (query, positive example, negative example, short text set);

[0062] S3.2. For the short texts in the generated (query, positive example, negative example, short text set), the quality control index M is used to quantitatively evaluate the quality of the short texts. The calculation formula is as follows:

[0063]

[0064] Where Z is the short text set corresponding to the positive example A, Ai and Aj are two different short texts in the short text set Z; sim(A,Ai) represents the similarity between the positive example A and the generated short text Ai, which is the cosine similarity of the two vectors calculated after encoding the two texts A and Ai into the vector space using the primary retrieval model obtained by S2; M i represents the quality control indicator of the short text Ai;

[0065] S3.3. Set the hyperparameters K=3 and R=0. For all short text sets, the quality control index of each short text data must be greater than 0. At most the top 3 short texts with the largest quality control index are retained to form an enhanced data set in the format of (query, positive example, negative example, enhanced short text set).

[0066] S4, based on the primary retrieval model, the data in the enhanced data set are aggregated to obtain a guidance vector, and the guidance vector is used as auxiliary information to train the primary retrieval model to obtain a final retrieval model; the specific process is:

[0067] S4.1. First, load the primary retrieval model obtained in S2, and divide the enhanced data set (query, positive example, negative example, enhanced short text set) obtained in S3 into several training batches with a batch size of 32. The training task is therefore divided into several batches. In a batch of training tasks, the query, positive example, negative example, and enhanced short text set in the batch are all encoded as high-dimensional vectors, and the enhanced short text set does not pass through the pooling layer. Add all the short text high-dimensional vectors from the same query to obtain a concatenated vector; then pass the concatenated vector through a fully connected layer and a residual connection layer, and project it into a high-level semantic space; the concatenated vector of each query is split according to the vector position, and the knn algorithm is used to cluster and generate several phrase vectors based on the Euclidean distance between the vectors at each position; for each query vector, calculate the KL divergence between the query vector and each phrase vector, and use the phrase vector corresponding to the minimum KL divergence as the guide vector; the calculation formula of the guide vector is:

[0068] P=min u∈U (D KL (Q||u))

[0069] Among them, P represents the guidance vector, Q represents the high-dimensional vector encoded by the query, U represents the set of phrase vectors, u is a single phrase vector, and D KL represents the divergence function;

[0070] S4.2. Use the loss function L to perform training and learning to obtain the final retrieval model; the loss function L = L1 + L2, respectively as follows:

[0071] For the query and short text set, mark any short text corresponding to the query as a positive example and the rest as negative examples to calculate the cross entropy as the loss function L1;

[0072] L1=Σ-ylpgsim(Q,S)-(1-y)log(1-sim(Q,S))

[0073] Among them, sim represents the cosine distance between vectors; Q represents the high-dimensional vector encoded by the query; S represents the high-dimensional vector encoded by the short text; y is the actual label, when the short text is a positive example, y = 1, when the short text is a negative example, y = 0;

[0074] For the guidance vector and the positive example, the cross entropy between the guidance vector corresponding to any query and the positive example corresponding to the query and the negative example corresponding to the query is calculated as the loss function L2;

[0075] L2=Σ-ylogsim(P,D)-(1-y)log(1-sim(P,D))

[0076] Among them, sim represents the cosine distance of the vector; P represents the guidance vector; D represents the vector encoded by the positive example or the negative example; y is the actual label, when D is the vector encoded by the positive example, y=1, when D is the vector encoded by the negative example, y=0.

[0077] This embodiment conducts experiments on the paragraph retrieval dataset Multi-CPR, and MRR@10 is used as the evaluation indicator. Multi-CPR is a comprehensive dataset designed for multi-domain Chinese article retrieval. This dataset covers three very different fields: e-commerce, entertainment video, and medical care, and aims to provide research resources for cross-domain information retrieval; the data subset of each field contains millions of text paragraphs. In addition, this embodiment also conducts experiments on the dataset T2Ranking, which is a dataset constructed for large-scale Chinese article retrieval tasks. The dataset brings together more than 300,000 user queries and more than two million unique paragraph texts from real-world search engines. The method of the present invention is compared with the QL, BM25, DEw / BM25Neg, DEw / MinedNeg, and DPTDR algorithms, and the experimental results are shown in Tables 1 and 2.

[0078] Among them, QL (query likelihood) is a representative statistical language model that measures the relevance of paragraphs by modeling the generation of queries. BM25 is a widely used sparse retrieval baseline. DEw / BM25Neg is equivalent to DPR, which is the first work to use a pre-trained language model as the backbone for article retrieval tasks. DEw / MinedNeg, as described in ANCE, enhances the performance of DPR by globally sampling hard negative examples from the entire corpus. DPTDR is a work that uses on-the-fly tuning for dense vector retrieval.

[0079] Table 1 Experimental results of T2Ranking dataset

[0080]

[0081] Table 2 Comparison of MRR@10 indicators in Multi-CPR dataset experiments

[0082]

[0083] According to the experimental results, the data enhancement strategy implemented by the present invention significantly optimizes the performance of the model in terms of the performance indicator of average reciprocal ranking. This performance improvement can be attributed to three aspects: first, the use of data enhancement allows the model to perform more difficult tasks, thereby stimulating the model's learning ability; second, the reduction in text length after the data enhancement proposed by the present invention allows more training samples to be placed in the same video memory space, which alleviates the problem of inconsistent model training prediction targets; third, by introducing guidance vectors, this method effectively integrates diversified information from the same source, significantly reducing the difficulty of the model directly learning difficult data.

[0084] The present invention proposes a text data enhancement method for improving the performance of vector retrieval for retrieval tasks. The starting point of this method is to decouple the data with the help of a large language model to generate better quality training data, and to screen the data through quality evaluation indicators; through a two-stage training method, especially using the construction process of the guidance vector to make full use of the enhanced data, the performance of the vector retrieval model is finally improved.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A text data enhancement method for improving vector retrieval performance, characterized in that: The following steps are involved: S1, obtaining a data set and preprocessing the data set; S2, using the preprocessed data set to perform comparative learning training on the pre-trained model to obtain a trained primary retrieval model; S3, semantically decoupling the positive example documents in the preprocessed data set based on the large language model, decomposing the positive example documents into several short texts; and filtering the short text data to obtain an enhanced data set; S4. Aggregate the data in the enhanced data set based on the primary retrieval model to construct a guidance vector, and use the guidance vector to guide the training of the primary retrieval model to obtain a final retrieval model.

2. A text data enhancement method for improving vector retrieval performance according to claim 1, characterized in that: In S1, the data set includes a query file, a query-positive example correspondence file, and a document collection file.

3. A text data enhancement method for improving vector retrieval performance according to claim 2, characterized in that: In S1, the data set is preprocessed, specifically: If the data set contains a correspondence file between queries and negative examples, the positive and negative documents are determined based on the correspondence files between the queries and the positive and negative examples, and the data set format is processed into (query, positive, negative); if the data set lacks a correspondence file between queries and negative examples, the retrieval model is used to retrieve similar documents in the entire document collection as query negative examples, and the positive documents are determined based on the correspondence file between the queries and the positive examples, and the data set format is processed into (query, positive, negative).

4. A text data enhancement method for improving vector retrieval performance according to claim 3, characterized in that: In S2, the training process of the primary retrieval model is as follows: S2.1, use the pre-trained model to initialize its pre-trained checkpoint, the pre-trained model can be used as an encoder to map all queries and positive examples in the data set obtained in S1 to the same vector space; S2.2, divide the data set obtained in S1 into several training batches. In the training task of any batch, first encode the query, positive document and negative document in the batch into vectors at the same time; S2.

3. In the vector space of the same training batch, for any query vector, its corresponding negative examples include all documents except its own positive example document. The distances between the query vector and each positive example document vector and each negative example document vector are calculated, and the calculated distances are smoothed by the softmax function to make them distributed between (0,1). Finally, the cross entropy is used as the loss function for training to obtain a trained primary retrieval model. The loss function formula is as follows: in, is the result of the distance between the query vector and the positive document vector or the negative document vector after being processed by the softmax function; y is the actual label, when the query vector and the positive document vector calculate the distance y = 1, when the query vector and the negative document vector calculate the distance, y = 0.

5. A text data enhancement method for improving vector retrieval performance according to claim 3 or 4, characterized in that: The specific process of S3 is as follows: S3.

1. For a training set of the data set in S1, construct each positive example document into an instruction using a prompt template and input it into the large language model to obtain the semantically decoupled short text output of each positive example document, and then change the training set format to (query, positive example, negative example, short text set); S3.

2. For the short texts in the generated (query, positive example, negative example, short text set), the quality control index M is used to quantitatively evaluate the quality of the short texts. The calculation formula is as follows: M i =Σ Aj∈Z yes(A,Ai)-∑ Ai,Aj∈Z yes(Ai,Aj) Wherein, Z is the short text set corresponding to the positive example A, Ai and Aj are two different short texts in the short text set Z; sim(A,Ai) represents the similarity between the positive example A and the generated short text Ai, which is the cosine similarity of the two vectors calculated after encoding the two texts A and Ai into the vector space using the primary retrieval model; M i represents the quality control indicator of the short text Ai; S3.

3. Set the hyperparameters K and R. For all short text sets, the quality control index of each short text data must be greater than R. At most K short texts with the largest quality control index are retained to form an enhanced data set in the format of (query, positive example, negative example, enhanced short text set).

6. A text data enhancement method for improving vector retrieval performance according to claim 5, characterized in that: The specific process of S4 is as follows: S4.

1. First, the primary retrieval model is loaded, and the enhanced data set is divided into several training batches. In a training task of a batch, the query, positive example, negative example, and enhanced short text set in the batch are all encoded into high-dimensional vectors, wherein the enhanced short text set does not pass through the pooling layer. The concatenated vector is obtained by adding the high-dimensional vectors of all short texts from the same query. The concatenated vector is then passed through a fully connected layer and a residual connection layer, and projected into a high-level semantic space. The concatenated vector of each query is split according to the vector position, and clustered based on the Euclidean distance between the vectors at each position to generate several phrase vectors. For each query vector, the KL divergence between the query vector and each phrase vector is calculated, and the phrase vector corresponding to the minimum KL divergence is used as the guide vector. The calculation formula of the guide vector is: P=min u∈U (D KL (Q||u)) Among them, P represents the guidance vector, Q represents the high-dimensional vector encoded by the query, U represents the set of phrase vectors, u is a single phrase vector, and D KL represents the divergence function; S4.

2. Use the loss function L to perform training and learning to obtain the final retrieval model; the loss function L = L1 + L2, respectively as follows: For the query and short text set, mark any short text corresponding to the query as a positive example and the rest as negative examples to calculate the cross entropy as the loss function L1; L1=Σ-ylogsim(Q,S)-(1-y)log(1-sim(Q,S)) Among them, sim represents the cosine distance between vectors; Q represents the high-dimensional vector encoded by the query; S represents the high-dimensional vector encoded by the short text; y is the actual label, when the short text is a positive example, y = 1, when the short text is a negative example, y = 0; For the guidance vector and the positive example, the cross entropy between the guidance vector corresponding to any query and the positive example corresponding to the query and the negative example corresponding to the query is calculated as the loss function L2; L2=∑-ylogsim(P,D)-(1-y)log(1-sim(P,D)) Among them, sim represents the cosine distance of the vector; P represents the guidance vector; D represents the vector encoded by the positive example or the negative example; y is the actual label, when D is the vector encoded by the positive example, y=1, when D is the vector encoded by the negative example, y=0.

7. A text data enhancement method for improving vector retrieval performance according to claim 1, characterized in that: The pre-trained model includes RetroMAE or GTE.

8. A text data enhancement method for improving vector retrieval performance according to claim 1, characterized in that: The large language model includes GPT-4 or a general meaning model.

Citation Information

Patent Citations

  • Text retrieval method based on multi-view comparative learning

    CN114880452A

  • Text retrieval matching model training method and device, electronic equipment and medium

    CN115146021A

  • Conversational information retrieval method based on pre-training language model

    CN115391500A

  • Retrieval enhancement generation method based on vector similarity matching optimization

    CN117573815A

  • Unsupervised pre-training method and system for strengthening vector retrieval capability

    CN118014045A

Cited By

  • Intelligent question and answer method and system based on dual-stage retrieval and generation

    CN122045379A