Processing method and device for retrieving document embedding characteristics

By introducing an embedded feature enhancement model into the RAG framework, integrating the correlation characteristics of the search document and input text and adding isolators, the problem of interference in the generation quality in the traditional RAG framework is solved, and the generation quality and identification capabilities of large language models are improved.

CN120387426APending Publication Date: 2025-07-29BEIJING DP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510521382.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When the traditional RAG framework uses the search document supplementary model to input text, it fails to effectively utilize the correlation and isolation characteristics of the search document and the input text, resulting in the generation quality being interfered with by irrelevant or redundant information.

Method used

An embedded feature enhancement model is designed, and the correlation features between the document and the model input text are searched through fusion, and isolating embedding vectors are added to each document embedding feature to form an enhanced RAG framework for training of multi-class text generation tasks.

Benefits of technology

It improves the ability of large language models to identify search documents, reduces redundant information interference, and improves the ability to ensure the quality of generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387426A_ABST
    Figure CN120387426A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and a device for processing embedded features of a retrieved document. The method comprises the following steps of: setting a processing model for performing correlation and isolator feature addition on the embedded features of the retrieved document as an embedded feature enhancement model; adding the embedded feature enhancement model into a traditional RAG framework to obtain a corresponding enhanced RAG framework; and training the embedded feature enhancement model based on the multi-class text generation tasks by taking the enhanced RAG framework as a text generation task execution main body. Through the embedded feature enhancement model, the identification capability of the large language model on the retrieval document can be improved, redundant information interference is reduced, and the guarantee capability on the generation quality of the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method and apparatus for processing embedded features of retrieved documents. Background Art

[0002] The Retrieval Augmented Generation (RAG) framework is a technical framework for improving the quality of text generation by large language models (LLMs). Generally, a retrieval system / model and LLMs are configured under the RAG framework. The retrieval system / model is used for external document queries, and LLMs are used for processing text generation tasks. The working process of the traditional RAG framework is roughly as follows: The text planned to be input into LLMs this time is used as the current input text and encoded based on the embedding encoding rules of LLMs. At the same time, based on the built-in retrieval system / model, relevant documents are retrieved from the resource document library using the current input text as a query, and all retrieved documents are encoded based on the embedding encoding rules of LLMs; then all retrieved documents are used as supplements to the current input text, and the embedding encodings of the current input text and all retrieved documents are spliced; finally, the spliced vector is fed into LLMs for text generation task processing.

[0003] In principle, the RAG framework can enhance the understanding, reasoning, and generation capabilities of LLMs by using the retrieved documents output by the retrieval system / model as supplementary knowledge. However, when the traditional RAG framework uses retrieved documents to supplement the content of the model input text, it only supplements knowledge based on a simple text splicing method, and does not use features such as the relevance between the retrieved documents and the model input text, and the relevance between the retrieved documents as supplementary information and inputs them into LLMs. This may instead cause LLMs to be interfered by irrelevant or redundant information when generating answers, reducing the generation quality. Summary of the Invention

[0004] The object of the present invention is to provide a processing method, device, electronic device and computer-readable storage medium for retrieving document embedding features in view of the defects of the prior art. The present invention designs a processing model, namely an embedding feature enhancement model, for adding correlation and delimiter features to the embedding features of retrieved documents; adds the embedding feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework; and trains the embedding feature enhancement model based on multiple types of text generation tasks (such as question-answering tasks, translation tasks, etc.) with the enhanced RAG framework as the main body for executing the text generation task. On the one hand, the embedding feature enhancement model provided by the present invention fuses the correlation features reflecting the correlation between the retrieved document and the model input text and between the retrieved documents with the text embedding features, improving the semantic richness of the document embedding features; on the other hand, a standard delimiter embedding vector is added to each document embedding feature, improving the recognition ability of the LLMs for the document embedding features. Through the embedding feature enhancement model of the present invention, the recognition ability of the LLMs for the retrieved documents can be improved, the interference of redundant information to the LLMs can be reduced, and the guarantee ability for the generation quality of the LLMs can be improved.

[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a method for processing retrieved document embedding features, the method comprising:

[0006] Set up a processing model for adding correlation and delimiter features to the embedding features of retrieved documents, denoted as an embedding feature enhancement model; and add the embedding feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework;

[0007] Train the embedding feature enhancement model based on multiple types of text generation tasks with the enhanced RAG framework as the main body for executing the text generation task; the multiple types of text generation tasks at least include question-answering tasks and translation tasks.

[0008] Preferably, both the traditional RAG framework and the enhanced RAG framework are used to perform text generation processing on the basis of the currently input input text T in and output the corresponding generated text T out ;

[0009] The traditional RAG framework at least includes an input embedding encoding layer, a retrieval model, a document embedding encoding layer, an embedding splicing layer and a large language model; the large language model is implemented based on the Transformer architecture;

[0010] The connection relationships of the components of the traditional RAG framework are as follows: the input end of the input embedding and encoding layer is connected to the input end of the traditional RAG framework, and the output end is connected to the first input end of the embedding splicing layer; the input end of the retrieval model is connected to the input end of the traditional RAG framework, and the output end is connected to the input end of the document embedding and encoding layer. The retrieval model is also connected to the document library outside the framework; the output end of the document embedding and encoding layer is connected to the second input end of the embedding splicing layer; the output end of the embedding splicing layer is connected to the input end of the large language model; the output end of the large language model is connected to the input end of the traditional RAG framework;

[0011] The processing process of the text generation task of the traditional RAG framework is as follows: the input embedding and encoding layer performs text embedding and encoding on the input text T in according to the embedding and encoding rules of the large language model to obtain the corresponding embedding vector x q and sends it to the embedding splicing layer; and the retrieval model uses the input text T in as the current query text to perform document retrieval on the document library to obtain the corresponding retrieved document sequence D and sends it to the document embedding and encoding layer; and the document embedding and encoding layer performs text embedding and encoding on the retrieved document sequence D according to the embedding and encoding rules of the large language model to obtain the corresponding embedding tensor X D and sends it to the embedding splicing layer; and the embedding splicing layer performs vector splicing processing on the embedding vector x q and the embedding tensor X D to obtain the corresponding input tensor X and sends it to the large language model; and the large language model performs text generation processing according to the input tensor X to obtain the corresponding generated text T out and outputs it;

[0012] The enhanced RAG framework includes the input embedding and encoding layer, the retrieval model, the document embedding and encoding layer, the embedding feature enhancement model, the embedding splicing layer, and the large language model;

[0013] The connection relationships of the components of the enhanced RAG framework are as follows: The input end of the input embedding encoding layer is connected to the input end of the enhanced RAG framework, and the output end is respectively connected to the first input end of the embedding splicing layer and the first input end of the embedding feature enhancement model; The input end of the retrieval model is connected to the input end of the enhanced RAG framework, and the output end is connected to the input end of the document embedding encoding layer. The retrieval model is also connected to the document library outside the framework; The output end of the document embedding encoding layer is connected to the second input end of the embedding feature enhancement model; The output end of the embedding feature enhancement model is connected to the second input end of the embedding splicing layer; The output end of the embedding splicing layer is connected to the input end of the large language model; The output end of the large language model is connected to the input end of the enhanced RAG framework;

[0014] The processing process of the text generation task of the enhanced RAG framework is as follows: The input embedding encoding layer performs text embedding encoding on the input text T according to the embedding encoding rules of the large language model in to obtain the corresponding embedding vector x q and send it to the embedding splicing layer and the embedding feature enhancement model; And the retrieval model uses the input text T in as the current query text to perform document retrieval on the document library to obtain the corresponding retrieved document sequence D and send it to the document embedding encoding layer; And the document embedding encoding layer performs text embedding encoding on the retrieved document sequence D according to the embedding encoding rules of the large language model to obtain the corresponding embedding tensor X D and send it to the embedding feature enhancement model; And the embedding feature enhancement model adds relevance and separator features to the retrieved document embedding features according to the embedding vector x q and the embedding tensor X D to obtain the corresponding enhanced embedding tensor E D and send it to the embedding splicing layer; And the embedding splicing layer performs vector splicing processing on the embedding vector x q and the enhanced embedding tensor E D to obtain the corresponding input tensor X' and send it to the large language model; And the large language model performs text generation processing according to the input tensor X' to obtain the corresponding generated text T out and outputs it;

[0015] The shape of the embedding vector x q is L q ×K; L q is the total number of tokenized tokens corresponding to the input text T in , and K is the preset token embedding feature dimension;

[0016] The retrieved document sequence D includes multiple retrieved documents d i ; 1 ≤ index i ≤ N D , N D being the total number of retrieved documents in the currently described retrieved document sequence D;

[0017] The embedding tensor X D has a shape of being the total number of tokenized tokens corresponding to the retrieved document d i ; the embedding tensor X D is composed of N D embedding vectors , and each of the embedding vectors has a shape of the corresponding The embedding vector corresponds one-to-one with the retrieved document d i ;

[0018] The shape of the input tensor X is The input tensor X is sequentially concatenated by the embedding vector x q and N D of the embedding vectors ;

[0019] The enhanced embedding tensor E D has a shape of The enhanced embedding tensor E D is composed of N D enhanced embedding vectors , and the enhanced embedding vector corresponds one-to-one with the retrieved document d i ; each of the enhanced embedding vectors has a shape of the corresponding Each of the enhanced embedding vectors is sequentially concatenated by the separator embedding vector s, the corresponding embedding vector and the correlation feature vector ; the separator embedding vector s and the correlation feature vector both have a shape of 1×K; the N D separator embedding vectors s of the enhanced embedding tensor E D have consistent vector data;

[0020] The shape of the input tensor X' is The input tensor X' is sequentially concatenated by the embedding vector x q and N D of the enhanced embedding vectors ;

[0021] Preferably, the embedding feature enhancement model is used to generate, according to the currently input embedding vector x q and the embedding tensor X D to increase the relevance and separator features in the retrieved document embedding features to obtain the corresponding enhanced embedding tensor E D ;

[0022] The embedding feature enhancement model includes a dense encoder, a relevance feature recognition layer, a dimensionality increase mapping layer, a positional embedding layer, a self-attention encoder, an embedding mapping layer, a relevance embedding fusion layer, and a separator embedding layer; the dense encoder is implemented based on a pre-trained BERT series model; the dimensionality increase mapping layer and the embedding mapping layer are each implemented based on a linear network; the self-attention encoder is implemented based on the Encoder model of the Transformer architecture;

[0023] The first and second input ends of the dense encoder are respectively connected to the first and second input ends of the embedding feature enhancement model, and the output end is connected to the input end of the relevance feature recognition layer; the output end of the relevance feature recognition layer is connected to the input end of the dimensionality increase mapping layer; the output end of the dimensionality increase mapping layer is connected to the input end of the positional embedding layer; the output end of the positional embedding layer is connected to the input end of the self-attention encoder; the output end of the self-attention encoder is connected to the input end of the embedding mapping layer; the output end of the embedding mapping layer is connected to the first input end of the relevance embedding fusion layer; the second input end of the relevance embedding fusion layer is connected to the second input end of the embedding feature enhancement model, and the output end is connected to the input end of the separator embedding layer; the output end of the separator embedding layer is connected to the output end of the embedding feature enhancement model;

[0024] The dense encoder is used to perform high-dimensional feature encoding on the embedding vector x q to obtain the corresponding feature vector v q ; and perform high-dimensional feature encoding on each of the embedding vectors D of the embedding tensor X to obtain the corresponding feature vectors and the N D obtained feature vectors form the corresponding feature tensor V D ; and send the feature vector v q and the feature tensor V D to the relevance feature recognition layer; the feature vector v q , have the same feature dimension;

[0025] The relevance feature recognition layer is used to process each of the feature vectors With the feature vector v q The cosine vector similarity is calculated to obtain the corresponding similarity And for each of the said feature vectors The cosine vector similarity with its previous feature vector is calculated to obtain the corresponding similarity And the said similarity corresponding to the feature vector is set as the preset default similarity; and for each of the said feature vectors The cosine vector similarity with its next feature vector is calculated to obtain the corresponding similarity And the said similarity corresponding to the feature vector is set as the said default similarity; and from each of the said feature vectors The said similarity corresponding to is formed into a corresponding correlation feature vector The said similarity corresponding to is formed into a corresponding correlation feature vector And the N D obtained correlation feature vectors are formed into the corresponding correlation feature tensor C D and sent to the dimensionality-raising mapping layer;

[0026] The calculation method of the said similarity is:

[0027] The calculation method of the said similarity is:

[0028] The calculation method of the said similarity is:

[0029] cosine() is the cosine vector similarity calculation function;

[0030] The shape of the said correlation feature vector is 1×3; the shape of the correlation feature tensor C D is N D ×3;

[0031] The dimensionality-raising mapping layer is used to raise the feature dimension of the correlation feature tensor C D to the feature dimension N of the self-attention encoder EN to obtain the corresponding dimensionality-raising tensor Y D and send it to the position embedding layer;

[0032] The calculation method of the dimensionality-raising mapping layer is: Y D = CD W1 + B1;

[0033] The shape of the weight matrix W1 is 3×N EN , and the shape of the offset matrix B1 is 1×N EN ;

[0034] The upsampled tensor Y D consists of N D upsampled vectors ; The shape of the upsampled vector is 1×N EN ; The shape of the upsampled tensor Y D is N D ×N EN ;

[0035] The position embedding layer is used to set a position embedding vector with a feature dimension matching the feature dimension N for each of the upsampled vectors of the upsampled tensor Y according to the position embedding coding rule of the self-attention encoder D and form the corresponding position embedding tensor P from the obtained N position embedding vectors EN ; And the upsampled tensor Y and the position embedding tensor P D are added to obtain the corresponding embedding tensor Z which is sent to the self-attention encoder; The embedding tensor Z D consists of N D embedding vectors D , Z D = Y D + P D ; The shape of the embedding vector D is 1×N D , and the shape of the embedding tensor Z D is N The embedding vector has a shape of 1×N EN , and the shape of the embedding tensor Z D is N D ×N EN ;

[0036] The self-attention encoder is used to perform self-attention encoding processing on the embedding tensor Z D and output the corresponding feature tensor H D which is sent to the embedding mapping layer; The feature tensor H D consists of N D feature vectors , and the shape of the feature vector is 1×N EN , and the shape of the feature tensor H D is ND ×N EN ;

[0037] The embedding mapping layer is used to map the feature dimension of the feature tensor H D to the token embedding feature dimension K through a linear network mapping to obtain a corresponding mapping tensor M D and send it to the correlation embedding fusion layer;

[0038] The calculation method of the embedding mapping layer is: M D = H D W2 + B2;

[0039] The shape of the weight matrix W2 is N EN ×K, and the shape of the offset matrix B1 is 1×K;

[0040] The mapping tensor M D consists of N D mapping vectors ; The shape of the mapping vector is 1×K; The shape of the mapping tensor M D is N D ×K;

[0041] The correlation embedding fusion layer is used to fuse the embedding tensor X D and the mapping tensor M D by vector concatenation to obtain a corresponding fusion tensor R D and send it to the separator embedding layer; The fusion tensor R D consists of N D fusion vectors ; The fusion vector is sequentially concatenated by the corresponding embedding vector and the mapping vector ; The shape of the fusion vector is The shape of the fusion tensor R D is

[0042] The separator embedding layer is used to perform embedding encoding on a preset separator token according to the embedding encoding rules of the large language model to obtain a corresponding separator embedding vector s; and the separator embedding vector s and each fusion vector of the fusion tensor R D are sequentially concatenated to form a corresponding enhanced embedding vector ; and N obtained enhanced embedding vectors D form a corresponding enhanced embedding tensor E ​D and output; the enhanced embedding vector has a shape of the enhanced embedding tensor E D has a shape of

[0043] Preferably, training the embedding feature enhancement model based on multiple types of text generation tasks with the enhanced RAG framework as the text generation task execution subject specifically includes:

[0044] Constructing a model data set by collecting a large amount of data on the input text and labeled output text of the RAG framework for the multiple types of text generation tasks, denoted as the corresponding first data set; and training the embedding feature enhancement model based on the first data set with the enhanced RAG framework as the text generation task execution subject;

[0045] Among them, the first data set includes multiple first data records; the first data record includes a first training text, a first labeled text, and a first label probability tensor; the first training text is the input text of the RAG framework for a type of text generation task, the first labeled text is the labeled output text corresponding to the current first training text; when the text generation task type corresponding to the first training text is a question-answering task, the current first training text is the question text, and the current first labeled text is the corresponding standard answer text; when the text generation task type corresponding to the first training text is a translation task, the current first training text is the pre-translation text, and the first labeled text is the corresponding post-translation text; the first label probability tensor is composed of multiple first label probability vectors, the total number of the first label probability vectors matches the total number of text segmentations of the first labeled text, and the first label probability vectors correspond one by one to the text segmentations of the first labeled text; the first label probability vector is the probability distribution vector of the current text segmentation in the vocabulary vector space of the large language model, the vector length of the first label probability vector is consistent with the feature dimension of the vocabulary vector space, each vector data is a segmentation probability, corresponding one by one to the segmentations in the model vocabulary of the large language model, and only the vector data corresponding to the current text segmentation in the first label probability vector is 1, and the rest are all 0; the feature dimension of the vocabulary vector space is consistent with the total number of segmentations in the model vocabulary; all text segmentations of the first labeled text come from the model vocabulary.

[0046] Furthermore, training the embedding feature enhancement model based on the first data set with the enhanced RAG framework as the text generation task execution subject specifically includes:

[0047] Step 51, divide the first data set into two sub-data sets according to a preset first splitting ratio, denoted as the corresponding first training set and first evaluation set; and form the corresponding current training parameter set from the model parameters of the dimension-increasing mapping layer, self-attention encoder, and embedding mapping layer of the embedding feature enhancement model;

[0048] Among them, both the first training set and the first evaluation set include multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first splitting ratio;

[0049] Step 52, take each first data record in the first training set as the corresponding current training record; and take the first training text of the current training record as the corresponding input text T in Input it into the enhanced RAG framework for processing, and extract the text probability distribution tensor generated by the large language model during this processing as the corresponding first prediction probability tensor; and form a corresponding first prediction-label pair from the first prediction probability tensor and the first label probability tensor of the current training record; out

[0050] Step 53, input all the obtained first prediction-label pairs into a preset first model loss function for calculation to obtain the corresponding first loss value;

[0051] Among them, the first model loss function is implemented based on the cross-entropy loss function or the mean square error loss function;

[0052] Step 54, identify whether the first loss value satisfies a preset first loss value range; if it satisfies, go to Step 55; if it does not satisfy, perform one round of parameter modulation on the current training parameter set of the embedding feature enhancement model in the direction that makes the first model loss function reach the minimum value based on a preset first model optimizer, and return to Step 52 to continue training at the end of this round of parameter modulation;

[0053] Among them, the first model optimizer includes at least the Adam optimizer and the SGD optimizer;

[0054] Step 55, take each first data record in the first evaluation set as the corresponding current evaluation record; and take the first training text of the current evaluation record as the corresponding input text T in Input it into the enhanced RAG framework for processing, and extract the text probability distribution tensor generated by the large language model during this processing as the corresponding first prediction probability tensor; and form a corresponding first prediction-label pair from the first prediction probability tensor and the first label probability tensor of the current evaluation record; out ​Extract the text probability distribution tensor as the corresponding second prediction probability tensor; and form a corresponding second prediction-label pair from the second prediction probability tensor and the first label probability tensor of the current training record; and bring all the obtained second prediction-label pairs into a preset first model evaluation function for calculation to obtain a corresponding first evaluation value;

[0055] Among them, the first model evaluation function is implemented based on the MAE function, MSE function or RMSE function;

[0056] Step 56, identify whether the first evaluation value satisfies a preset first evaluation value range; if not, return to step 52 to continue training; if so, stop training and confirm that the model training of the embedding feature enhancement model is completed.

[0057] A second aspect of the embodiments of the present invention provides an apparatus for implementing the processing method for retrieving document embedding features described in the first aspect above. The apparatus includes: a model design module and a model training module;

[0058] The model design module is used to set a processing model for adding correlation and delimiter features to the embedding features of the retrieved document, denoted as the embedding feature enhancement model; and add the embedding feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework;

[0059] The model training module is used to train the embedding feature enhancement model based on multiple types of text generation tasks with the enhanced RAG framework as the text generation task execution body; the multiple types of text generation tasks at least include a question-answering task and a translation task.

[0060] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor and a transceiver;

[0061] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0062] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0063] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0064] Embodiments of the present invention provide a method, apparatus, electronic device, and computer-readable storage medium for processing embedded features of retrieved documents. As can be seen from the above, embodiments of the present invention design a processing model, namely an embedded feature enhancement model, for adding correlation and delimiter features to the embedded features of retrieved documents; add the embedded feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework; and use the enhanced RAG framework as the execution entity for text generation tasks to train the embedded feature enhancement model based on various types of text generation tasks (such as question-answering tasks, translation tasks, etc.). On the one hand, the embedded feature enhancement model provided by the embodiments of the present invention fuses the correlation features reflecting the association between the retrieved document and the model input text and between the retrieved documents with the text embedding features, improving the semantic richness of the document embedding features; on the other hand, a standard delimiter embedding vector is added to each document embedding feature, improving the recognition ability of LLMs for the document embedding features. Through the embodiments of the present invention, the recognition ability of LLMs for retrieved documents can be improved, and the interference of redundant information on LLMs can be reduced. While further improving the generation quality of LLMs, the quality guarantee ability is also effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 Schematic diagram of a method for processing embedded features of retrieved documents provided in Embodiment 1 of the present invention;

[0066] Figure 2 Schematic diagram of the framework structure of the traditional RAG framework provided in Embodiment 1 of the present invention;

[0067] Figure 3 Schematic diagram of the framework structure of the enhanced RAG framework provided in Embodiment 1 of the present invention;

[0068] Figure 4 Schematic diagram of the model structure of the embedded feature enhancement model provided in Embodiment 1 of the present invention;

[0069] Figure 5 Module structure diagram of a device for processing embedded features of retrieved documents provided in Embodiment 2 of the present invention;

[0070] Figure 6 Schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] Embodiment 1 of the present invention provides a method for processing the embedded features of a retrieved document. As Figure 1 shown in the schematic diagram of the method for processing the embedded features of a retrieved document provided in Embodiment 1 of the present invention, the method mainly includes the following steps:

[0073] Step 1: Set a processing model for adding relevance and delimiter features to the embedded features of the retrieved document, denoted as the embedded feature enhancement model; and add the embedded feature enhancement model to the traditional RAG framework to obtain the corresponding enhanced RAG framework.

[0074] Here, the traditional RAG framework and the enhanced RAG framework in the embodiments of the present invention are both used to perform text generation processing on the basis of the currently input input text T in and output the corresponding generated text T out .

[0075] As Figure 2 shown in the schematic diagram of the framework structure of the traditional RAG framework provided in Embodiment 1 of the present invention, the traditional RAG framework in the embodiments of the present invention at least includes an input embedding encoding layer, a retrieval model, a document embedding encoding layer, an embedding splicing layer, and a large language model.

[0076] It should be noted that the retrieval model in the framework is a preset retrieval system or model, which can be a software / hardware retrieval software, module, device, equipment, server, system, or platform with a fixed retrieval logic; it can also be a retrieval model implemented based on another deep learning model or large language model. The purpose of this retrieval model is to perform relevant document retrieval on the resource information library, that is, the document library outside the framework, according to the given query and output the corresponding retrieved document sequence. It should also be noted that the large language model in the framework is a model implemented based on the Transformer architecture.

[0077] The connection relationship of the components of the traditional RAG framework is as follows: the input end of the input embedding encoding layer is connected to the input end of the traditional RAG framework, and the output end is connected to the first input end of the embedding splicing layer; the input end of the retrieval model is connected to the input end of the traditional RAG framework, and the output end is connected to the input end of the document embedding encoding layer. The retrieval model is also connected to the document library outside the framework; the output end of the document embedding encoding layer is connected to the second input end of the embedding splicing layer; the output end of the embedding splicing layer is connected to the input end of the large language model; the output end of the large language model is connected to the input end of the traditional RAG framework.

[0078] The processing process of the text generation task of the traditional RAG framework is as follows: the input embedding encoding layer performs text embedding encoding on the input text T in according to the embedding encoding rules of the large language model to obtain the corresponding embedding vector x qSend it to the embedding splicing layer; and the retrieval model uses the input text T in As the current query text, perform document retrieval on the document library to obtain the corresponding retrieved document sequence D, and send it to the document embedding encoding layer; and the document embedding encoding layer performs text embedding encoding on the retrieved document sequence D according to the embedding encoding rules of the large language model to obtain the corresponding embedding tensor X D Send it to the embedding splicing layer; and the embedding splicing layer processes the embedding vector x q And the embedding tensor X D Perform vector splicing processing to obtain the corresponding input tensor X and send it to the large language model; and the large language model performs text generation processing according to the input tensor X to obtain the corresponding generated text T out And output it.

[0079] As Figure 3 As shown in the framework structure diagram of the enhanced RAG framework provided in Embodiment 1 of the present invention, the enhanced RAG framework of the present invention includes an input embedding encoding layer, a retrieval model, a document embedding encoding layer, an embedding feature enhancement model, an embedding splicing layer, and a large language model. Except for the embedding feature enhancement model in the enhanced RAG framework, the remaining input embedding encoding layer, retrieval model, document embedding encoding layer, embedding feature enhancement model, embedding splicing layer, and large language model all come from the traditional RAG framework.

[0080] The connection relationships of the components of the enhanced RAG framework are as follows: the input end of the input embedding encoding layer is connected to the input end of the enhanced RAG framework, and the output ends are respectively connected to the first input end of the embedding splicing layer and the first input end of the embedding feature enhancement model; the input end of the retrieval model is connected to the input end of the enhanced RAG framework, and the output end is connected to the input end of the document embedding encoding layer. The retrieval model is also connected to the document library outside the framework; the output end of the document embedding encoding layer is connected to the second input end of the embedding feature enhancement model; the output end of the embedding feature enhancement model is connected to the second input end of the embedding splicing layer; the output end of the embedding splicing layer is connected to the input end of the large language model; the output end of the large language model is connected to the input end of the enhanced RAG framework.

[0081] The processing process of the text generation task of the enhanced RAG framework is as follows: the input embedding encoding layer performs text embedding encoding on the input text T according to the embedding encoding rules of the large language model in To obtain the corresponding embedding vector x q Send it to the embedding splicing layer and the embedding feature enhancement model; and the retrieval model uses the input text T in As the current query text, perform document retrieval on the document library to obtain the corresponding retrieved document sequence D, and send it to the document embedding encoding layer; and the document embedding encoding layer performs text embedding encoding on the retrieved document sequence D according to the embedding encoding rules of the large language model to obtain the corresponding embedding tensor X DSend to the embedded feature enhancement model; and the embedded feature enhancement model is based on the embedding vector x q and the embedding tensor X D Increase the relevance and delimiter features in the retrieved document embedding features to obtain the corresponding enhanced embedding tensor E D Send to the embedding concatenation layer; and the embedding concatenation layer performs vector concatenation processing on the embedding vector x q and the enhanced embedding tensor E D Perform vector concatenation processing to obtain the corresponding input tensor X', send it to the large language model; and the large language model performs text generation processing based on the input tensor X' to obtain the corresponding generated text T out And output.

[0082] The data structures of various types of data mentioned in the above two frameworks are as follows:

[0083] 1) The embedding vector x q has a shape of L q ×K; L q is the total number of tokenized tokens corresponding to the input text T in , and K is the preset token embedding feature dimension;

[0084] Here, the embedding vector x q is the vector data obtained after the input embedding encoding layer in the traditional / enhanced RAG framework performs embedding encoding on the input text T in ; the encoding process of the input embedding encoding layer is implemented based on the embedding encoding rules of the large language model, and the embedding encoding rules of the large language model usually first tokenize the input text T in based on the model vocabulary to obtain a tokenized sequence, then perform token conversion on the tokens to obtain a tokenized token sequence, then perform embedding encoding on each token in the sequence to obtain a token embedding vector with a feature dimension of K, and then form the embedding vector x from all the token embedding vectors q ; because the text length and tokenization status of each input text T in are different, the total number of tokenized tokens L q is a variable and is determined by the text length of each input text T in and the corresponding tokenization status; however, the token embedding feature dimension K is a fixed constant, and this token embedding feature dimension K is also the feature dimension of the input embedding vector of the large language model in the traditional / enhanced RAG framework. When the large language model does not change, this K is naturally a fixed constant;

[0085] 2) The retrieved document sequence D includes multiple retrieved documents d i ; 1 ≤ index i ≤ N D , N D is the total number of retrieved documents in the current retrieved document sequence D;

[0086] Here, N D can be a variable or a constant, depending on the retrieval function of the retrieval model in the traditional / enhanced RAG framework; if the retrieval model can dynamically adjust the number of output documents each time, then N D is a variable, determined by the total number of documents output by the retrieval model each time; if the number of documents output by the retrieval model is fixed each time, then N D is a constant, matching the fixed output quantity of the retrieval model;

[0087] 3) Embedded tensor X D consists of N D embedded vectors ; the embedded vectors correspond one-to-one with the retrieved document d i ; the shape of each embedded vector is the corresponding which is the total number of tokenized tokens corresponding to the retrieved document d i ; the shape of the embedded tensor X D is

[0088] Here, the embedded tensor X D is the tensor data obtained after the document embedding encoding layer in the traditional / enhanced RAG framework embeds and encodes the retrieved document sequence D; the encoding process of the document embedding encoding layer is also implemented based on the embedding encoding rules of the large language model, similar to the embedded vector x q ; for each retrieved document d i the corresponding embedded vector naturally has the shape of the corresponding And the embedded tensor X D can be regarded as being composed of all the embedded vectors stitched together, so the shape of the embedded tensor X D naturally is

[0089] 4) The shape of the input tensor X is The input tensor X is composed of the embedded vector x q and N D embedded vectors stitched together in sequence;

[0090] Here, the input tensor X is the input vector of the large language model generated by the traditional RAG framework using the retrieved document sequence D to supplement knowledge for the input text T in , which is essentially equivalent to simply concatenating the input text T in and the retrieved document sequence D and performing embedding encoding on the concatenated text to obtain the embedded tensor;

[0091] 5) Enhance the embedding tensor E D Consisting of N D enhanced embedding vectors The enhanced embedding vectors are in one-to-one correspondence with the retrieved document d i For each enhanced embedding vector the shape is the corresponding For each enhanced embedding vector is composed of a separator embedding vector s, the corresponding embedding vector and a correlation feature vector concatenated in sequence; the separator embedding vector s and the correlation feature vector both have a shape of 1×K; for the N D separator embedding vectors s of the enhanced embedding tensor E D the vector data are consistent; the shape of the enhanced embedding tensor E D is

[0092] Here, the enhanced embedding tensor E D is the embedding tensor output by the embedding feature enhancement model in the enhanced RAG framework. This embedding tensor not only includes the traditional retrieved document embedding vector, i.e., the embedding vector but also includes the correlation feature vector and the separator embedding vector s; the role of the separator embedding vector s is equivalent to adding a separator recognizable by the large language model when concatenating the input text T in and the retrieved document sequence D for each retrieved document d i which improves the independent recognition of the features of the retrieved document d i by the large language model, i.e., the enhanced embedding vector ; the role of the correlation feature vector is equivalent to incorporating into the embedding features of each retrieved document d i correlation features that can reflect the relationship between the retrieved document d i and the input text T in as well as the correlation between the retrieved document d i and the surrounding documents. Based on this feature, it helps the large language model to dynamically adjust the attention to the retrieved document d i thus improving the model's anti-interference ability while enhancing the model generation quality;

[0093] 6) The shape of the input tensor X’ is The input tensor X’ is composed of the embedding vector x q and N D enhanced embedding vectors concatenated in sequence;

[0094] Here, the input tensor X' is the input vector of the large language model generated after the enhanced RAG framework supplements knowledge to the input text T using the retrieved document sequence D in for knowledge supplementation

[0095] As Figure 4 shown in the model structure diagram of the embedding feature enhancement model provided in Embodiment 1 of the present invention, the newly added embedding feature enhancement model in the enhanced RAG framework of the present invention is used to add correlation and separator features to the retrieved document embedding features according to the currently input embedding vector x q and the embedding tensor X D to obtain the corresponding enhanced embedding tensor E D .

[0096] The embedding feature enhancement model of the embodiment of the present invention includes: a dense encoder, a correlation feature recognition layer, a dimensionality increase mapping layer, a position embedding layer, a self-attention encoder, an embedding mapping layer, a correlation embedding fusion layer, and a separator embedding layer. It should be noted that the dense encoder is implemented based on a pre-trained BERT series model, and the model parameters of the dense encoder may not be modulated in subsequent model training. It should also be noted that the dimensionality increase mapping layer and the embedding mapping layer of the embodiment of the present invention are each implemented based on a linear network. It should also be noted that the self-attention encoder of the embodiment of the present invention is implemented based on the Encoder model of the Transformer architecture, and bidirectional self-attention encoding is achieved through the multi-head attention module of the Encoder model

[0097] The connection relationships of the components of the embedding feature enhancement model are as follows: the first and second input ends of the dense encoder are respectively connected to the first and second input ends of the embedding feature enhancement model, and the output end is connected to the input end of the correlation feature recognition layer; the output end of the correlation feature recognition layer is connected to the input end of the dimensionality increase mapping layer; the output end of the dimensionality increase mapping layer is connected to the input end of the position embedding layer; the output end of the position embedding layer is connected to the input end of the self-attention encoder; the output end of the self-attention encoder is connected to the input end of the embedding mapping layer; the output end of the embedding mapping layer is connected to the first input end of the correlation embedding fusion layer; the second input end of the correlation embedding fusion layer is connected to the second input end of the embedding feature enhancement model, and the output end is connected to the input end of the separator embedding layer; the output end of the separator embedding layer is connected to the output end of the embedding feature enhancement model

[0098] The functions of the components of the embedding feature enhancement model are as follows

[0099] 1) Dense encoder:

[0100] The dense encoder of the embodiment of the present invention is used to perform high-dimensional feature encoding on the embedding vector x q to obtain the corresponding feature vector vq ; and perform high - dimensional feature encoding on each embedding vector D of the embedded tensor X to obtain corresponding feature vectors and form a corresponding feature tensor V D from the obtained N feature vectors D ; and send the feature vector v q and the feature tensor V D to the correlation feature recognition layer.

[0101] Here, the feature dimensions of the feature vectors v q , are the same.

[0102] 2) Correlation feature recognition layer:

[0103] The correlation feature recognition layer in the embodiment of the present invention is used to calculate the cosine vector similarity between each feature vector and the feature vector v q to obtain the corresponding similarity and calculate the cosine vector similarity between each feature vector and its previous feature vector to obtain the corresponding similarity and set the similarity corresponding to the feature vector as a preset default similarity; and calculate the cosine vector similarity between each feature vector and its subsequent feature vector to obtain the corresponding similarity and set the similarity corresponding to the feature vector as the default similarity; and form a corresponding correlation feature vector from the similarities corresponding to each feature vector and form a corresponding correlation feature tensor C D from the obtained N correlation feature vectors D and send it to the dimensionality - increasing mapping layer.

[0104] Here, the calculation method of the similarity is:

[0105]

[0106] where cosine() is the cosine vector similarity calculation function; the default similarity is a preset constant with a value between 0 and 1, which can be set according to application requirements, such as 0, 1, 0.5, etc.; the correlation feature vector The shape is 1×3; the correlation feature tensor C D has a shape of N D ×3.

[0107] 3) Dimensionality-raising mapping layer:

[0108] The dimensionality-raising mapping layer in the embodiment of the present invention is used to raise the feature dimension of the correlation feature tensor C D to the feature dimension N of the self-attention encoder EN to obtain the corresponding dimensionality-raising tensor Y D and send it to the position embedding layer.

[0109] Here, the calculation method of the dimensionality-raising mapping layer is:

[0110] Y D = C D W1 + B1;

[0111] Among them, the weight matrix W1 has a shape of 3×N EN , and the offset matrix B1 has a shape of 1×N EN ; the dimensionality-raising tensor Y D consists of N D dimensionality-raising vectors ; the dimensionality-raising vector has a shape of 1×N EN ; the dimensionality-raising tensor Y D has a shape of N D ×N EN .

[0112] 4) Position embedding layer:

[0113] The position embedding layer in the embodiment of the present invention is used to set a position embedding vector with a feature dimension matching the feature dimension N for each dimensionality-raising vector D of the dimensionality-raising tensor Y and form the corresponding position embedding tensor P EN from the obtained N position embedding vectors D ; and add the dimensionality-raising tensor Y and the position embedding tensor P D to obtain the corresponding embedding tensor Z D and send it to the self-attention encoder. D Here, the embedding tensor Z in the embodiment of the present invention D consists of N

[0114] embedding vectors D ; Z D = Y + P D = Y D + PD , The embedding vector has a shape of 1×N EN , and the embedding tensor Z D has a shape of N D ×N EN . The position embedding encoding rule of the self-attention encoder in the embodiments of the present invention is the same as that of the Transformer architecture.

[0115] 5) Self-attention encoder:

[0116] The self-attention encoder in the embodiments of the present invention is used to perform self-attention encoding processing on the embedding tensor Z D and output the corresponding feature tensor H D to be sent to the embedding mapping layer.

[0117] Here, the feature tensor H in the embodiments of the present invention D consists of N D feature vectors , and the feature vector has a shape of 1×N EN , and the feature tensor H D has a shape of N D ×N EN .

[0118] 6) Embedding mapping layer:

[0119] The embedding mapping layer in the embodiments of the present invention is used to map the feature dimension of the feature tensor H D to the token embedding feature dimension K through a linear network mapping to obtain the corresponding mapping tensor M D to be sent to the correlation embedding fusion layer.

[0120] Here, the calculation method of the embedding mapping layer is:

[0121] M D = H D W2 + B2;

[0122] where the weight matrix W2 has a shape of N EN ×K, and the offset matrix B1 has a shape of 1×K; the mapping tensor M D consists of N D mapping vectors , the mapping vector has a shape of 1×K, and the mapping tensor M D has a shape of N D ×K.

[0123] 7) Correlation embedding fusion layer:

[0124] The relevance embedding fusion layer of the embodiment of the present invention is used to fuse the embedding tensor X D and the mapping tensor M D to obtain the corresponding fusion tensor R D and send it to the separator embedding layer.

[0125] Here, the fusion tensor R D is composed of N D fusion vectors ; the fusion vector is sequentially spliced by the corresponding embedding vector and the mapping vector ; the shape of the fusion vector is The shape of the fusion tensor R D is

[0126] 8) Separator embedding layer:

[0127] The separator embedding layer of the embodiment of the present invention is used to perform embedding encoding on a preset separator token according to the embedding encoding rules of the large language model to obtain the corresponding separator embedding vector s; and the separator embedding vector s and each fusion vector of the fusion tensor R D are sequentially spliced to form the corresponding enhanced embedding vector And the obtained N D enhanced embedding vectors are used to form the corresponding enhanced embedding tensor E D and output.

[0128] Here, the preset separator token is the tokenized token corresponding to a pre-specified separator, and the separator is a separator recognizable by the large language model of the embodiment of the present invention. In addition, the shape of the enhanced embedding vector is The shape of the enhanced embedding tensor E D is

[0129] Step 2: Use the enhanced RAG framework as the main body of the text generation task to train the embedding feature enhancement model based on multiple types of text generation tasks;

[0130] Among them, the multiple types of text generation tasks at least include question answering tasks and translation tasks;

[0131] Specifically, it includes: Step 21, construct a model data set denoted as the corresponding first data set by collecting a large amount of data of the input text and the labeled output text of the RAG framework for multiple types of text generation tasks;

[0132] ​Among them, the first data set includes multiple first data records; the first data record includes a first training text, a first label text, and a first label probability tensor;

[0133] The first training text is the input text of the RAG framework for a type of text generation task, and the first label text is the label output text corresponding to the current first training text; when the text generation task type corresponding to the first training text is a question-and-answer task, the current first training text is the question text, and the current first label text is the corresponding standard answer text; when the text generation task type corresponding to the first training text is a translation task, the current first training text is the text before translation, and the first label text is the corresponding text after translation;

[0134] The first label probability tensor is composed of multiple first label probability vectors. The total number of first label probability vectors matches the total number of text segmentations of the first label text, and the first label probability vectors correspond one by one to the text segmentations of the first label text; the first label probability vector is the probability distribution vector of the current text segmentation in the vocabulary vector space of the large language model. The vector length of the first label probability vector is consistent with the feature dimension of the vocabulary vector space. Each vector data is a segmentation probability and corresponds one by one to the segmentation in the model vocabulary of the large language model. Only the vector data corresponding to the current text segmentation in the first label probability vector is 1, and the rest are all 0; the feature dimension of the vocabulary vector space is consistent with the total number of segmentations in the model vocabulary; all text segmentations of the first label text come from the model vocabulary;

[0135] Step 22, and use the enhanced RAG framework as the execution entity of the text generation task to train the embedding feature enhancement model based on the first data set;

[0136] Specifically, it includes: Step 221, divide the first data set into two sub-data sets according to a preset first segmentation ratio, denoted as the corresponding first training set and first evaluation set; and form the corresponding current training parameter set from the model parameters of the dimensionality increase mapping layer, self-attention encoder, and embedding mapping layer of the embedding feature enhancement model;

[0137] Here, the first segmentation ratio is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set include multiple first data records; the record total ratio of the first training set and the first evaluation set satisfies the first segmentation ratio;

[0138] Step 222, use each first data record in the first training set as the corresponding current training record; and use the first training text of the current training record as the corresponding input text T in Input it into the enhanced RAG framework for processing, and use the large language model in this processing process to generate the generated text T outExtract the text probability distribution tensor as the corresponding first prediction probability tensor; and form a corresponding first prediction-label pair from the first prediction probability tensor and the first label probability tensor of the current training record;

[0139] Here, a brief description of the processing flow of the large language model according to the embodiments of the present invention before generating the output text T out Before that, a simple explanation is given: Currently, all large language models implemented based on the Transformer architecture will obtain a text probability distribution tensor through a series of encoding and / or decoding processes before generating the output text T out Before that, a text probability distribution tensor is obtained through a series of encoding and / or decoding processes. The text probability distribution tensor is composed of one or more text probability distribution vectors. Each text probability distribution vector corresponds to an output word / token. The vector length of each text probability distribution vector matches the total number of tokenizations of the model vocabulary, that is, it is usually said to be consistent with the feature dimension of the vocabulary vector space. Each vector data of the text probability distribution vector corresponds to a token in the model vocabulary, and is the prediction probability of the output word / token corresponding to the current vector for the current vocabulary tokenization; when using the first training text of the current training record as the corresponding input text T in When inputting and enhancing the RAG framework for processing, only the process data of the large language model in the framework needs to be cached, and the generated text T can be extracted from the cached data out The corresponding text probability distribution tensor, that is, the first prediction probability tensor;

[0140] Step 223, bring all the obtained first prediction-label pairs into a preset first model loss function for calculation to obtain a corresponding first loss value;

[0141] Here, the first model loss function according to the embodiments of the present invention is implemented based on the cross-entropy loss function or the mean square error loss function;

[0142] Step 224, identify whether the first loss value meets a preset first loss value range; if it meets, go to step 225; if it does not meet, perform a round of parameter modulation on the current training parameter set of the embedding feature enhancement model based on a preset first model optimizer in the direction of minimizing the first model loss function, and return to step 222 to continue training at the end of this round of parameter modulation;

[0143] Here, the first loss value range according to the embodiments of the present invention is a preset numerical range; the first model optimizer includes at least the Adam optimizer and the SGD optimizer;

[0144] Step 225, use each first data record in the first evaluation set as the corresponding current evaluation record; and use the first training text of the current evaluation record as the corresponding input text T inProcessed by the input-enhanced RAG framework, and in this processing process, the large language model is used to generate the generated text T out Extract the text probability distribution tensor of and use it as the corresponding second prediction probability tensor; and form a corresponding second prediction-label pair from the second prediction probability tensor and the first label probability tensor of the current training record; and bring all the obtained second prediction-label pairs into the preset first model evaluation function for calculation to obtain the corresponding first evaluation value;

[0145] Here, the first model evaluation function of the embodiment of the present invention is implemented based on the MAE function, the MSE function or the RMSE function;

[0146] Step 226, identify whether the first evaluation value satisfies the preset first evaluation value range; if not, return to step 222 to continue training; if so, stop training and confirm that the model training of the embedding feature enhancement model ends.

[0147] Here, the first evaluation value range of the embodiment of the present invention is a preset numerical range.

[0148] After the model training ends, use the enhanced RAG framework to replace the traditional RAG framework to process the text generation task. In this way, during each text generation task processing of the enhanced RAG framework, the embedding feature enhancement model can perform embedding feature enhancement processing on the retrieved documents generated during the current process, and the subsequent embedding splicing layer can then send the input tensor X' with the enhanced embedding tensor E D to the large language model for text generation processing.

[0149] Figure 5 It is a module structure diagram of a processing device for the embedding features of retrieved documents provided by the second embodiment of the present invention. This device is a terminal device or a server for implementing the foregoing method embodiments, or can be a device that enables the foregoing terminal device or server to implement the foregoing method embodiments. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 5 shown, this device includes: a model design module 201 and a model training module 202.

[0150] The model design module 201 is used to set a processing model for adding relevance and delimiter features to the embedding features of retrieved documents, denoted as the embedding feature enhancement model; and add the embedding feature enhancement model to the traditional RAG framework to obtain the corresponding enhanced RAG framework.

[0151] The model training module 202 is used to train the embedding feature enhancement model based on multiple types of text generation tasks with the enhanced RAG framework as the main body of the text generation task; the multiple types of text generation tasks at least include question and answer tasks and translation tasks.

[0152] The processing device for retrieving document embedding features provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0153] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in hardware. For example, the model design module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0154] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0155] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the above computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0156] Figure 6 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 6 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0157] In Figure 6The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement the communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0158] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0159] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0160] An embodiment of the present invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing embedded features of retrieved documents. As can be seen from the above, an embodiment of the present invention designs a processing model for adding correlation and delimiter features to the embedded features of retrieved documents, namely, an embedded feature enhancement model; and adds the embedded feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework; and uses the enhanced RAG framework as the execution subject of the text generation task to train the embedded feature enhancement model based on multiple types of text generation tasks (such as question-answering tasks, translation tasks, etc.). On the one hand, the embedded feature enhancement model provided by the embodiment of the present invention fuses the correlation features reflecting the association degree between the retrieved document and the model input text and between the retrieved documents with the text embedded features, improving the semantic richness of the document embedded features; on the other hand, a standard delimiter embedding vector is added to each document embedded feature, improving the recognition ability of LLMs for the document embedded features. Through the embodiment of the present invention, the recognition ability of LLMs for retrieved documents can be improved, and the interference of redundant information on LLMs can be reduced. While further improving the generation quality of LLMs, the quality assurance ability is also effectively improved.

[0161] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be disposed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art.

[0162] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A processing method for retrieving document embedding features, characterized in that The method includes: Setting up a processing model for adding relevance and delimiter features to the embedded features of retrieved documents, denoted as the embedded feature enhancement model; and adding the embedded feature enhancement model to the traditional RAG framework to obtain the corresponding enhanced RAG framework; Using the enhanced RAG framework as the execution entity for text generation tasks to train the embedded feature enhancement model based on multiple types of text generation tasks; the multiple types of text generation tasks at least include question-answering tasks and translation tasks.

2. The method for processing the embedded features of retrieved documents according to claim 1, wherein The traditional RAG framework and the enhanced RAG framework are both used to input the text T according to the current input. in Perform text generation processing and output the corresponding generated text T out ; The traditional RAG framework at least includes an input embedding encoding layer, a retrieval model, a document embedding encoding layer, an embedding splicing layer, and a large language model; the large language model is implemented based on the Transformer architecture; The connection relationships of the components of the traditional RAG framework are as follows: the input end of the input embedding encoding layer is connected to the input end of the traditional RAG framework, and the output end is connected to the first input end of the embedding splicing layer; the input end of the retrieval model is connected to the input end of the traditional RAG framework, and the output end is connected to the input end of the document embedding encoding layer. The retrieval model is also connected to a document library outside the framework; the output end of the document embedding encoding layer is connected to the second input end of the embedding splicing layer; the output end of the embedding splicing layer is connected to the input end of the large language model; the output end of the large language model is connected to the input end of the traditional RAG framework; The processing process of the text generation task of the traditional RAG framework is as follows: the input embedding encoding layer performs text embedding encoding on the input text T according to the embedding encoding rules of the large language model in to obtain the corresponding embedding vector x q and send it to the embedding splicing layer; and the retrieval model uses the input text T in as the current query text to perform document retrieval on the document library to obtain the corresponding retrieved document sequence D and send it to the document embedding encoding layer; and the document embedding encoding layer performs text embedding encoding on the retrieved document sequence D according to the embedding encoding rules of the large language model to obtain the corresponding embedding tensor X D and send it to the embedding splicing layer; and the embedding splicing layer performs vector splicing processing on the embedding vector x q and the embedding tensor X D to obtain the corresponding input tensor X and send it to the large language model; and the large language model performs text generation processing according to the input tensor X to obtain the corresponding generated text T out and output it; The enhanced RAG framework includes the input embedding encoding layer, the retrieval model, the document embedding encoding layer, the embedded feature enhancement model, the embedding splicing layer, and the large language model; The connection relationships of the components of the enhanced RAG framework are as follows: the input end of the input embedding encoding layer is connected to the input end of the enhanced RAG framework, and the output end is respectively connected to the first input end of the embedding splicing layer and the first input end of the embedded feature enhancement model; the input end of the retrieval model is connected to the input end of the enhanced RAG framework, and the output end is connected to the input end of the document embedding encoding layer. The retrieval model is also connected to the document library outside the framework; the output end of the document embedding encoding layer is connected to the second input end of the embedded feature enhancement model; the output end of the embedded feature enhancement model is connected to the second input end of the embedding splicing layer; the output end of the embedding splicing layer is connected to the input end of the large language model; the output end of the large language model is connected to the input end of the enhanced RAG framework; The processing procedure for the text generation task of the enhanced RAG framework is as follows: The input embedding encoding layer performs text embedding encoding on the input text T according to the embedding encoding rules of the large language model in to obtain the corresponding embedding vector x q and send it to the embedding concatenation layer and the embedding feature enhancement model; and the retrieval model uses the input text T in as the current query text to perform document retrieval on the document library to obtain the corresponding retrieved document sequence D and send it to the document embedding encoding layer; and the document embedding encoding layer performs text embedding encoding on the retrieved document sequence D according to the embedding encoding rules of the large language model to obtain the corresponding embedding tensor X D and send it to the embedding feature enhancement model; and the embedding feature enhancement model adds relevance and separator features to the retrieved document embedding features according to the embedding vector x q and the embedding tensor X D to obtain the corresponding enhanced embedding tensor E D and send it to the embedding concatenation layer; and the embedding concatenation layer performs vector concatenation processing on the embedding vector x q and the enhanced embedding tensor E D to obtain the corresponding input tensor X ’ and send it to the large language model; and the large language model performs text generation processing according to the input tensor X ’ to obtain the corresponding generated text T out and output it; Embedding vector x q has a shape of L q ×K; L q is the total number of tokenized tokens corresponding to the input text T in and K is the preset token embedding feature dimension; The retrieved document sequence D includes multiple retrieved documents d i ; 1 ≤ index i ≤ N D , N D is the total number of retrieved documents in the currently described retrieved document sequence D; The embedded tensor X D has a shape of which is the total number of tokenized tokens corresponding to the retrieved document d i ; the embedded tensor X D consists of N D embedded vectors , and each of the embedded vectors has a shape of the corresponding The embedded vector corresponds one-to-one with the retrieved document d i ; The shape of the input tensor X is The input tensor X is composed of the embedding vector x q and N D of these embedding vectors concatenated in sequence; The enhanced embedding tensor E D has a shape of The enhanced embedding tensor E D is composed of N D enhanced embedding vectors which are in one-to-one correspondence with the retrieved document d ; Each of the enhanced embedding vectors i has a shape corresponding to Each of the enhanced embedding vectors is sequentially concatenated by a separator embedding vector s, the corresponding embedding vector and a relevance feature vector ; The separator embedding vector s and the relevance feature vector both have a shape of 1×K; The N separator embedding vectors s of the enhanced embedding tensor E D have consistent vector data; D ​ The input tensor X ’ has a shape of The input tensor X ’ is sequentially concatenated by the embedding vector x q and N D of the enhanced embedding vectors in sequence.

3. The method for processing the embedded features of retrieved documents according to claim 2, wherein The embedded feature enhancement model is used to, according to the currently input embedded vector x q and the embedded tensor X D increase the correlation and delimiter features in the retrieved document embedded features to obtain the corresponding enhanced embedded tensor E D ; The embedded feature enhancement model includes a dense encoder, a correlation feature recognition layer, a dimensionality elevation mapping layer, a position embedding layer, a self-attention encoder, an embedding mapping layer, a correlation embedding fusion layer, and a separator embedding layer; the dense encoder is implemented based on a pre-trained BERT series model; the dimensionality elevation mapping layer and the embedding mapping layer are each implemented based on a linear network; the self-attention encoder is implemented based on the Encoder model of the Transformer architecture; The first and second input ends of the dense encoder are respectively connected to the first and second input ends of the embedded feature enhancement model, and the output end is connected to the input end of the correlation feature recognition layer; the output end of the correlation feature recognition layer is connected to the input end of the dimensionality elevation mapping layer; the output end of the dimensionality elevation mapping layer is connected to the input end of the position embedding layer; the output end of the position embedding layer is connected to the input end of the self-attention encoder; the output end of the self-attention encoder is connected to the input end of the embedding mapping layer; the output end of the embedding mapping layer is connected to the first input end of the correlation embedding fusion layer; the second input end of the correlation embedding fusion layer is connected to the second input end of the embedded feature enhancement model, and the output end is connected to the input end of the separator embedding layer; the output end of the separator embedding layer is connected to the output end of the embedded feature enhancement model; The dense encoder is used to embed the vector x q Perform high-dimensional feature encoding to obtain the corresponding feature vector v q ; and for the embedding tensor X D Each of the embedding vectors Perform high-dimensional feature encoding to obtain the corresponding feature vector And the obtained N D The feature vector Composed of the corresponding feature tensor V D ; and the feature vector v q and the feature tensor V D Sending to the correlation feature recognition layer; The eigenvector v q and have the same feature dimension; The correlation feature recognition layer is used to process each of the feature vectors and the feature vector v q to calculate the cosine vector similarity therebetween to obtain the corresponding similarity and process each of the feature vectors and its previous feature vector to calculate the cosine vector similarity therebetween to obtain the corresponding similarity and set the similarity corresponding to the feature vector as a preset default similarity; and process each of the feature vectors and its next feature vector to calculate the cosine vector similarity therebetween to obtain the corresponding similarity and set the similarity corresponding to the feature vector as the default similarity; and form a corresponding correlation feature vector from the similarities corresponding to each of the feature vectors and form a corresponding correlation feature tensor C from the N obtained correlation feature vectors and send it to the dimensionality increase mapping layer and send it to the dimensionality increase mapping layer D The N obtained correlation feature vectors D are sent to the dimensionality increase mapping layer; The similarity is calculated as follows: The similarity is calculated as follows: The similarity is calculated as follows: cosine() is a cosine vector similarity calculation function; The correlation feature vector has a shape of 1×3; the correlation feature tensor C D has a shape of N D ×3; The dimensionality - raising mapping layer is used to raise the feature dimension of the correlation feature tensor C D to the feature dimension N of the self - attention encoder EN through a linear network mapping, and obtain the corresponding dimensionality - raised tensor Y D and send it to the position embedding layer; The calculation method of the dimensionality-raising mapping layer is: Y D = C D W1 + B1; The shape of the weight matrix W1 is 3×N EN , and the shape of the bias matrix B1 is 1×N EN ; The dimensionality - raised tensor Y D is composed of N D dimensionality - raised vectors ; the dimensionality - raised vector has a shape of 1×N EN ; the dimensionality - raised tensor Y D has a shape of N D ×N EN ; The position embedding layer is used to set a position embedding vector with a feature dimension matching the feature dimension N for each of the upsampled vectors of the upsampled tensor Y according to the position embedding encoding rule of the self-attention encoder D of the upsampled tensor Y and form a corresponding position embedding tensor P composed of the obtained N EN position embedding vectors; and send the corresponding embedding tensor Z obtained by adding the upsampled tensor Y and the position embedding tensor P D to the self-attention encoder; the embedding tensor Z is composed of N D embedding vectors, Z D = Y D + P D ; the shape of the embedding vector D is 1×N D and the shape of the embedding tensor Z is N D ×N D ; D , The shape of the embedding vector is 1×N EN , and the shape of the embedding tensor Z D is N D ×N EN ; The self-attention encoder is used to perform self-attention encoding processing on the embedding tensor Z D and output the corresponding feature tensor H D and send it to the embedding mapping layer; the feature tensor H D consists of N D feature vectors ; the shape of the feature vector is 1×N EN , and the shape of the feature tensor H D is N D ×N EN ; The embedding mapping layer is used to map the feature dimension of the feature tensor H D to the token embedding feature dimension K through a linear network mapping to obtain a corresponding mapping tensor M D and send it to the correlation embedding fusion layer; The calculation method of the embedding mapping layer is: M D = H D W2 + B2; The shape of the weight matrix W2 is N EN ×K, and the shape of the bias matrix B1 is 1×K; The mapping tensor M D consists of N D mapping vectors ; the shape of the mapping vector is 1×K; the shape of the mapping tensor M D is N D ×K; The correlation embedding fusion layer is used to fuse the embedding tensor X D and the mapping tensor M D to obtain the corresponding fusion tensor R D and send it to the separator embedding layer; the fusion tensor R D is composed of N D fusion vectors ; the fusion vector is sequentially concatenated by the corresponding embedding vector and the mapping vector ; the shape of the fusion vector is The shape of the fusion tensor R D is The separator embedding layer is used to perform embedding encoding on a preset separator token according to the embedding encoding rule of the large language model to obtain the corresponding separator embedding vector s; and the separator embedding vector s and each of the fusion vectors of the fusion tensor R D are sequentially concatenated to form the corresponding enhanced embedding vector and N obtained enhanced embedding vectors D are used to form the corresponding enhanced embedding tensor E and output; the shape of the enhanced embedding vector D is and the shape of the enhanced embedding tensor E is D ​ 4. The method for processing the retrieval document embedding feature according to claim 2, wherein Taking the enhanced RAG framework as the execution subject of the text generation task to train the embedded feature enhancement model based on multiple types of text generation tasks, specifically including: Constructing a model data set, denoted as the corresponding first data set, by collecting a large amount of data of the RAG framework input text and the label output text of the multiple types of text generation tasks; and taking the enhanced RAG framework as the execution subject of the text generation task to train the embedded feature enhancement model based on the first data set; Among them, the first data set includes a plurality of first data records; the first data records include first training texts, first label texts, and first label probability tensors; the first training texts are input texts of the RAG framework for a type of text generation task, and the first label texts are label output texts corresponding to the current first training texts; when the text generation task type corresponding to the first training text is a question-answering task, the current first training text is a question text, and the current first label text is the corresponding standard answer text; when the text generation task type corresponding to the first training text is a translation task, the current first training text is the text before translation, and the first label text is the corresponding text after translation; the first label probability tensor is composed of a plurality of first label probability vectors, and the total number of the first label probability vectors matches the total number of text segmentations of the first label text, and the first label probability vectors correspond one-to-one with the text segmentations of the first label text; the first label probability vector is a probability distribution vector of the current text segmentation in the vocabulary vector space of the large language model, and the vector length of the first label probability vector is consistent with the feature dimension of the vocabulary vector space, and each vector data is a segmentation probability and corresponds one-to-one with the segmentations in the model vocabulary of the large language model, and only the vector data corresponding to the current text segmentation in the first label probability vector is 1, and the rest are all 0; the feature dimension of the vocabulary vector space is consistent with the total number of segmentations in the model vocabulary; all text segmentations of the first label text come from the model vocabulary.

5. The processing method for retrieving document embedding features according to claim 4, characterized in that Taking the enhanced RAG framework as the execution subject of the text generation task to train the embedding feature enhancement model based on the first data set specifically includes: Step 51, divide the first data set into two sub-data sets denoted as the corresponding first training set and first evaluation set based on a preset first segmentation ratio; and form a corresponding current training parameter set from the model parameters of the dimensionality increase mapping layer, self-attention encoder, and embedding mapping layer of the embedding feature enhancement model; Among them, both the first training set and the first evaluation set include a plurality of the first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 52: Take each of the first data records in the first training set as the corresponding current training record; and take the first training text of the current training record as the corresponding input text T in Input it into the enhanced RAG framework for processing, and use the large language model in this processing to generate the generated text T out Extract the text probability distribution tensor used by the large language model during this processing as the corresponding first prediction probability tensor; and form a corresponding first prediction-label pair from the first prediction probability tensor and the first label probability tensor of the current training record Step 53, bring all the obtained first prediction-label pairs into a preset first model loss function for calculation to obtain a corresponding first loss value; Among them, the first model loss function is implemented based on a cross-entropy loss function or a mean square error loss function; Step 54, identify whether the first loss value meets a preset first loss value range; if it meets, go to Step 55; if it does not meet, perform a round of parameter modulation on the current training parameter set of the embedding feature enhancement model in the direction of minimizing the first model loss function based on a preset first model optimizer, and return to Step 52 to continue training at the end of this round of parameter modulation; Among them, the first model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 55, take each of the first data records in the first evaluation set as the corresponding current evaluation record; and take the first training text of the current evaluation record as the corresponding input text T in Input it into the enhanced RAG framework for processing, and use the large language model in this processing to generate the generated text T out Extract the text probability distribution tensor of the generated text T as the corresponding second prediction probability tensor; and form a corresponding second prediction-label pair with the second prediction probability tensor and the first label probability tensor of the current training record; and bring all the obtained second prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value; Among them, the first model evaluation function is implemented based on the MAE function, the MSE function, or the RMSE function; Step 56, identify whether the first evaluation value satisfies a preset first evaluation value range; if not, return to step 52 to continue training; if so, stop training and confirm that the model training of the embedding feature enhancement model ends.

6. An apparatus for performing the processing method of retrieving document embedding features according to any one of claims 1-5, characterized in that, The device includes: a model design module and a model training module; The model design module is used to set a processing model for adding correlation and delimiter features to the embedding features of the retrieved document, denoted as the embedding feature enhancement model; and add the embedding feature enhancement model to the traditional RAG framework to obtain a corresponding enhanced RAG framework; The model training module is used to train the embedding feature enhancement model based on multiple types of text generation tasks with the enhanced RAG framework as the main body for text generation tasks; the multiple types of text generation tasks at least include a question-and-answer task and a translation task.

7. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-5; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-5.