Enhanced retrieval generation method based on potential fusion LoRA

Through an enhanced search generation method based on potentially fusion LoRA, combined with mixed similarity scores, dynamic gating and low-rank adaptation components, the problems of dynamic knowledge screening efficiency and model stability in the existing technology are solved, and efficient and stable knowledge fusion and generation are achieved, suitable for applications such as intelligent question-and-answer and text summary.

CN120596624AActive Publication Date: 2025-09-05HANGZHOU ZHONGKE RUIJIAN TECH CO LTD

Patent Information

Application Number
CN202510690787.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing search generation methods have room for optimization in terms of dynamic knowledge screening efficiency and model stability, especially in the dynamic integration of multi-source knowledge in complex scenarios.

Method used

Using an enhanced search generation method based on potential fusion LoRA, through a three-level coupling architecture of the search layer, adaptation layer and output layer, hybrid similarity score, dynamic gating mechanism and low-rank adaptation components are used to achieve efficient knowledge fusion and model stability, including hybrid similarity score, dynamic gating method to filter high correlation vectors, low-rank adaptation component feature fusion and large language model parameter freezing.

Benefits of technology

The model's knowledge utilization efficiency, reasoning accuracy, and scenario applicability are improved, ensuring generation quality and reducing deployment costs. It is suitable for scenarios such as intelligent question answering and text summarization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596624A_ABST
    Figure CN120596624A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to an enhanced retrieval generation method based on potential fusion LoRA. The method comprises the following steps: realizing knowledge fusion and dynamic optimization through a three-level architecture of a retrieval layer, an adaptation layer and an output layer; the retrieval layer encodes an input text into a query vector by adopting a pre-training model, screens Top-n correlation vectors from a knowledge base based on mixed similarity and splices the Top-n correlation vectors; the adaptation layer filters low-correlation retrieval results through a dynamic gating mechanism, performs low-rank decomposition and feature fusion in combination with LoRA, reduces the calculation amount and optimizes the knowledge injection efficiency; the output layer freezes large language model parameters, retains the general generation capability to output texts, and avoids knowledge forgetting caused by fine tuning. According to the method, the problems of noise interference, high calculation overhead and knowledge forgetting in an existing method are solved through mixed similarity scoring, dynamic gating and LoRA collaborative design and a freezing strategy, and the method is suitable for a question and answer system and a text generation scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to an enhanced retrieval generation method based on potential fusion LoRA. Background Art

[0002] As natural language processing technology places increasing demands on knowledge integration and generation capabilities, the retrieval-augmented generation (RAG) method has become a mainstream technical path for improving the effectiveness of tasks such as question-answering systems and text generation by integrating external knowledge bases to enhance model reasoning capabilities. Existing research is mostly based on directly splicing retrieval results or fine-tuning all parameters based on pre-trained language models. Although it can inject domain knowledge in specific scenarios, there is still room for improvement in terms of dynamic knowledge screening efficiency and model stability.

[0003] The Chinese invention application with publication number CN119578557A discloses an LLM capability enhancement method with a hybrid enhancement strategy. By acquiring corpus data and external literature data in the target application field, denoising, word segmentation and semantic indexing are performed to establish a searchable heterogeneous knowledge base; then, the corpus data is classified using topic analysis and word vector clustering methods to form a basic corpus set; based on the basic corpus set, the large language model is updated with multiple rounds of parameters using a domain fine-tuning algorithm, and a fusion threshold between the new parameters and the original parameters is set; relevant document fragments are retrieved from the heterogeneous knowledge base through a matching algorithm, and semantic similarity evaluation is used to screen content that meets domain relevance, which is then fused with the fine-tuned model. The fusion results are processed using a context fusion algorithm and an entity linking method to complete syntactic analysis and generate accurate answers for the target application field.

[0004] Low-rank adaptation (LoRA) technology achieves lightweight model adaptation through parameter decomposition. Its efficiency has been verified in single-task fine-tuning scenarios, but it has not yet been deeply integrated with the retrieval generation framework. In this context, there is an urgent need for a technical solution that takes into account knowledge fusion accuracy, computational efficiency and model generalization capabilities to meet the needs of dynamic integration of multi-source knowledge in complex scenarios and promote the practical application of retrieval generation technology. Summary of the Invention

[0005] The purpose of the present invention is to address the problems existing in the background technology and propose an enhanced retrieval generation method based on potential fusion LoRA.

[0006] The technical solution of the present invention is an enhanced retrieval generation method based on potential fusion LoRA, comprising:

[0007] Retrieval layer: Encode the input text into a query vector Q and retrieve the top-n relevant vectors from the vector knowledge base through a hybrid similarity scoring method;

[0008] Adaptation layer: It uses dynamic gating to filter highly correlated vectors and uses the low-rank adaptation component LoRA for feature fusion. It generates residual features through the low-rank decomposition matrix and dynamically adjusts the weight distribution of the top-n search results in combination with the fully connected layer.

[0009] Output layer: The fused features output by the adaptation layer are input into a large language model with frozen parameters to generate the final text. The parameters of the large language model remain fixed during the training and inference stages.

[0010] Preferably, the hybrid similarity scoring method is:

[0011]

[0012] Among them, Score i It represents the comprehensive score that measures the matching degree between the query vector and the knowledge base vector; a represents the weight coefficient; K i Represents the vectorized representation of the i-th fragment in the knowledge base; cos(Q,K i ) represents the cosine similarity, which measures the difference between Q and K i Similarity in the vector direction, the value range is [-1,1]; ||QK i ||2 represents the Euclidean distance, that is, the distance between vector Q and K i Absolute distance in space.

[0013] Preferably, the activation conditions of the dynamic gating method are:

[0014]

[0015] Here, τ represents the hyperparameter threshold, i.e., the threshold τ of the dynamic gating mechanism; g(x) represents the dynamic gating coefficient; x represents the semantic information currently being processed; W represents the trainable weight matrix; and σ(·) represents the Sigmoid activation function.

[0016] Preferably, the threshold τ of the dynamic gating mechanism is dynamically adjusted by the ratio of positive and negative samples in the training dataset.

[0017] Preferably, the LoRA component implements feature dimensionality reduction and mapping through low-rank decomposition matrices A and B;

[0018] The feature fusion formula is:

[0019] h lora =h base +g(x)·LoRA([h base ;k1;...;k n ]);

[0020] Among them, h lora represents the hidden layer representation after fusion; h baserepresents the hidden layer output of the basic model; LoRA(·) represents the low-rank adaptation component; k1,…,k n Represents the retrieved knowledge fragment vector; [h base ;k1;...;k n ] represents the feature splicing operation.

[0021] Preferably, the training loss function of the adaptation layer includes:

[0022] LoRA loss function L lora , used to supervise the accuracy of knowledge fusion;

[0023] Gating loss function L gate , used to optimize the decision weights of dynamic gating.

[0024] Preferably, the LoRA loss function L lora is the cross entropy loss:

[0025]

[0026] Gating loss function L gate is the binary cross entropy loss:

[0027]

[0028] Where N represents the number of samples; y i Indicates the output logits score of the correct token; x i Represents the output logits score of the current token; c represents the binary label value.

[0029] Preferably, the knowledge base retrieval at the retrieval layer uses a hierarchical-navigable-small-world-graph (HNSW) algorithm to accelerate vector matching.

[0030] Preferably, in the training process of the adaptation layer, only the parameters of the dynamic gating network, the fully connected layer and the LoRA component are optimized, and the parameters of the embedding model of the retrieval layer and the large language model of the output layer are kept frozen.

[0031] Preferably, the weight distribution is implemented through a fully connected layer, specifically including: performing dimensionality reduction processing on the concatenated matrix of the Top-n search results, generating attention weights based on contextual semantics, and adjusting the contribution ratio of different knowledge sources.

[0032] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:

[0033] This paper designs an enhanced retrieval generation method based on potential fusion LoRA. Through three key technologies: gated LoRA update, dynamic weight allocation, and efficient knowledge injection, the following comprehensive technical effects are achieved: while maintaining a lightweight design, the model's knowledge utilization efficiency, reasoning accuracy, and scenario applicability are improved, providing an efficient and reliable solution for dynamic knowledge enhancement generation of large language models:

[0034] The search layer uses a hybrid similarity scoring formula that combines weighted calculations of cosine similarity and Euclidean distance to comprehensively evaluate the semantic relevance and spatial proximity between query vectors and knowledge base fragments, significantly improving the accuracy and robustness of search results.

[0035] The adaptation layer filters low-relevant search content through a dynamic gating mechanism and combines low-rank adaptation technology to decompose and fuse features. This reduces the number of parameters while retaining the ability to inject core knowledge, reducing ineffective calculations and improving training efficiency.

[0036] The output layer uses a large language model parameter freezing strategy to avoid the risk of knowledge overwriting during fine-tuning and ensure that the model maintains its general generation capabilities when integrating external knowledge.

[0037] While ensuring generation quality, this invention achieves efficient execution of retrieval generation tasks through hierarchical architecture design and lightweight adaptation mechanism. It is suitable for scenarios such as intelligent question answering and text summarization, and has significant deployment cost advantages and technical practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flowchart of an enhanced retrieval generation method based on potential fusion LoRA proposed in the present invention;

[0039] Reference symbols: Decoder Input: the initial input vector of the decoder, representing the original user text after preprocessing; A: low-rank matrix, which projects the input features into a low-dimensional latent space; B: low-rank matrix, which maps the low-dimensional features back to the original model dimension and forms a low-rank decomposition with matrix A; h lora : hidden layer representation after fusion; h base : The hidden layer output of the basic model; Top-n vector: The top n knowledge fragments most relevant to the input query Q retrieved from the vector knowledge base. DETAILED DESCRIPTION

[0040] Example 1, as Figure 1 As shown, the present invention proposes an enhanced retrieval generation method based on potential fusion LoRA, which constructs a three-level coupling architecture of retrieval layer, adaptation layer and output layer. Its specific implementation steps include:

[0041] Level 1, the retrieval layer: responsible for vectorizing the input text and searching the vector knowledge base, specifically:

[0042] Input text encoding, that is, the input text sequence is encoded into a query vector Q through the pre-trained Embedding model;

[0043] It should be noted that the core function of the Embedding model is to map high-dimensional, discrete text data (such as words, sentences, and documents) into low-dimensional, continuous real number vectors (i.e., embedding vectors).

[0044] Knowledge base retrieval: The query vector Q is retrieved from the vector knowledge base (e.g., Faiss uses HNSW (Hierarchical Navigable Small World graphs)) to obtain the top-n vectors. The top-n vectors are then re-ranked by similarity calculation, fusing cosine similarity and Euclidean distance, and weighted by the formula:

[0045]

[0046] Among them, Score i represents a comprehensive score that measures the degree of matching between the query vector and the knowledge base vector; a represents a weight coefficient, which is used to balance the contribution ratio of cosine similarity and Euclidean distance. In this embodiment, a is a hyperparameter and defaults to 0.7; K i Represents the vectorized representation of the i-th fragment in the knowledge base; cos(Q,K i ) represents the cosine similarity, which measures the difference between Q and K i Similarity in the vector direction, the value range is [-1,1]; ||QK i ||2 represents the Euclidean distance (L2 norm), that is, the distance between vector Q and K i The absolute distance in space, the smaller the value, the closer it is;

[0047] Vector concatenation: Retrieve the top-n related vectors with the highest cosine similarity to Q and perform tensor concatenation of the query vector Q and the retrieved top-n vectors according to the following rules:

[0048] Step 1. Construct a fusion matrix of dimension (n+1)×dim, where dim represents the vector dimension.

[0049] Step 2: Take Q as the first row vector and arrange the top-n vectors in descending order of similarity as the subsequent matrix;

[0050] Level 2, Adaptation Layer: Knowledge fusion optimization is achieved through gating mechanism, dynamic weight allocation and LoRA components. The adaptation layer includes dynamic gating coefficients, fully connected layers and LoRA components. The weight matrix of the dynamic gating coefficient g(x) is initialized using Xavier and L2 regularization is added. Specifically:

[0051] Dynamic gating screening, where the gating network activates or inhibits the knowledge fusion path based on the relevance of the input data:

[0052] Among them, τ represents the hyperparameter threshold, that is, the threshold of the dynamic gating mechanism, which is dynamically adjusted by the ratio of positive and negative samples in the training data set. If the current gating value is less than τ, the LoRA component of the potential representation fusion will not be called; g(x) represents the dynamic gating coefficient, whose core function is to determine whether to activate a specific module (LoRA component) through threshold judgment, thereby flexibly adjusting the knowledge injection strategy; x represents the input feature vector, which is usually the query vector Q or knowledge vector K output by the retrieval layer i , representing the semantic information currently being processed; W represents the trainable weight matrix, which performs a linear transformation on the input x to generate a scalar value for gating decisions; through training optimization, the knowledge relevance weights in different scenarios are learned; σ(·) represents the Sigmoid activation function;

[0053] It should be noted that LoRA (Low-Rank Adaptation) is a technology for fine-tuning models, which aims to reduce the training time and computing resource requirements of the model on specific tasks while maintaining the performance of the model. LoRA technology is implemented by decomposing the parameters of the model into the product of two low-rank matrices, which represent the weights of the model and the adaptation layer respectively. This method allows the model to adapt to new tasks or data sets through the adaptation layer while maintaining the low-rank structure of the original parameters. Specifically, LoRA decomposes the parameters of each layer into two smaller matrices (lower-rank matrices), which are multiplied and added to a smaller adaptation layer, thereby reducing the total number of parameters and computational complexity.

[0054] Dynamic weight allocation and dimensionality reduction, that is, the LoRA component module uses the fully connected layer to dynamically reduce the input fusion Q and the matrix of the top-n vector to the specified dimension, and completes the low-rank adaptation through the matrices A and B, outputs the residual update, and compares the retrieved vector related information with the original model input content h base Calculate and generate h lora , h lora By h base , gating and LoRA components to obtain:

[0055] h lora =h base +g(x)·LoRA([h base ;k1;...;k n ]);

[0056] Among them, h lora Represents the fused hidden layer representation, combined with the basic model output (h base ) and external knowledge (k1,…,k n ), enhanced features generated by LoRA components; h base represents the hidden layer output of the basic model; LoRA(·) represents the low-rank adaptation component (Low-Rank Adaptation); k1,…,k n Represents the retrieved knowledge fragment vector, which is obtained from the external knowledge base through a vector retrieval engine (such as Faiss); [h base ;k1;...;k n ] represents feature splicing operation;

[0057] Loss function, including LoRA loss L lora And the gated loss function:

[0058] LoRA loss function (cross entropy loss) L lora To ensure that the model correctly answers the questions in the dataset task:

[0059]

[0060] Where N represents the number of samples; y i Indicates the output logits score of the correct token; x i Indicates the output logits score of the current token;

[0061] Gated loss function (binary cross entropy) L gate , supervise the gating network decision to ensure that only highly relevant search content triggers knowledge fusion:

[0062] Among them, c represents the binary label value, that is, the current label is on or off;

[0063] Based on this: In the adaptation layer, the following adaptive processing is achieved through gating and fully connected layer networks:

[0064] (1) When the similarity between the Top-n vector and Q is detected to be lower than the preset threshold, the gating unit automatically suppresses the information transmission of the low-correlation vector;

[0065] (2) For highly correlated subset vectors, feature enhancement is performed using the attention weights learned by the fully connected layer;

[0066] (3) The mechanism is automatically optimized through back-propagation and finally outputs a weighted fusion feature representation;

[0067] Level 3 generates the final text based on the large language model with frozen parameters. This involves inputting the information processed by the adaptation layer into the large language model structure. Leveraging the large language model's existing powerful language understanding and generation capabilities, the final output is generated based on the previously integrated information without changing its original parameters. This ensures that the model maintains its own stable language processing paradigm while utilizing new knowledge. Specifically:

[0068] C1. Frozen large language model generation: The output layer uses a large language model with frozen parameters. The fused features output by the adaptation layer are input into the model to generate the final text. The freezing strategy preserves the general language capabilities of the model while avoiding the problem of knowledge forgetting caused by fine-tuning.

[0069] C2. The model training process is:

[0070] D1. Dataset construction: Label the correlation between the input question Q and the search results (positive and negative sample ratio 1:1) and construct the adaptation layer training data;

[0071] D2. Model construction: Freeze the large language model and the Embedding model, and add the gating network, fully connected layer, and LoRA components to the adaptation layer. The loss function of the gating network is L gate , the loss function of the fully connected layer and LoRA network is L lora ;

[0072] That is, only the gating network, fully connected layer, and LoRA module of the adaptation layer are optimized, and the parameters of the embedding model and the large language model are fixed;

[0073] D3, Iterative Optimization: The network model inputs text, obtains Q through the word embedding step, and uses the retriever to search the knowledge base to obtain the Top-n for splicing. The gated network is trained based on whether the labels in the dataset are associated. If the labels are associated, the subsequent fully connected layers and LoRA components are trained and backpropagated until convergence.

[0074] That is, if the retrieval result is related to the input (label is 1), the LoRA component is activated to update the parameters; otherwise, LoRA training is skipped to reduce invalid calculations.

[0075] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. An enhanced retrieval generation method based on potential fusion LoRA, characterized in that, include: Retrieval layer: Encode the input text into a query vector Q and retrieve the top-n relevant vectors from the vector knowledge base through a hybrid similarity scoring method; Adaptation layer: It uses dynamic gating to filter highly correlated vectors and uses the low-rank adaptation component LoRA for feature fusion. It generates residual features through the low-rank decomposition matrix and dynamically adjusts the weight distribution of the top-n search results in combination with the fully connected layer. Output layer: The fused features output by the adaptation layer are input into a large language model with frozen parameters to generate the final text. The parameters of the large language model remain fixed during the training and inference stages.

2. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: The hybrid similarity scoring method is: Among them, Score i It represents the comprehensive score that measures the matching degree between the query vector and the knowledge base vector; a represents the weight coefficient; K i Represents the vectorized representation of the i-th fragment in the knowledge base; cos(Q,K i ) represents the cosine similarity, which measures the difference between Q and K i Similarity in the vector direction, the value range is [-1,1]; ||QK i ||2 represents the Euclidean distance, that is, the distance between vector Q and K i Absolute distance in space.

3. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: The activation conditions of the dynamic gating method are: Here, τ represents the hyperparameter threshold, i.e., the threshold τ of the dynamic gating mechanism; g(x) represents the dynamic gating coefficient; x represents the semantic information currently being processed; W represents the trainable weight matrix; and σ(·) represents the Sigmoid activation function.

4. The enhanced search generation method based on potential fusion LoRA according to claim 3 is characterized in that: The threshold τ of the dynamic gating mechanism is dynamically adjusted according to the ratio of positive and negative samples in the training dataset.

5. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: The LoRA component achieves feature dimensionality reduction and mapping through low-rank decomposition matrices A and B; The feature fusion formula is: h lora =h base +g(x)·LoRA([h base ;k1;...;k n ]); Among them, h lora represents the hidden layer representation after fusion; h base represents the hidden layer output of the basic model; LoRA(·) represents the low-rank adaptation component; k1,…,k n Represents the retrieved knowledge fragment vector; [h base ;k1;...;k n ] represents the feature splicing operation.

6. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: The training loss function of the adaptation layer includes: LoRA loss function L lora , used to supervise the accuracy of knowledge fusion; Gating loss function L gate , used to optimize the decision weights of dynamic gating.

7. The enhanced search generation method based on potential fusion LoRA according to claim 6 is characterized in that: LoRA loss function L lora is the cross entropy loss: Gating loss function L gate is the binary cross entropy loss: Where N represents the number of samples; y i Indicates the output logits score of the correct token; x i Represents the output logits score of the current token; c represents the binary label value.

8. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: The knowledge base retrieval in the retrieval layer uses the Hierarchical-Navigable-Small-World-Graph (HNSW) algorithm to accelerate vector matching.

9. The enhanced search generation method based on potential fusion LoRA according to claim 1 is characterized in that: During the training process of the adaptation layer, only the parameters of the dynamic gating network, fully connected layer, and LoRA components are optimized. The parameters of the embedding model of the retrieval layer and the large language model of the output layer remain frozen.

10. The enhanced search generation method based on potential fusion LoRA according to claim 1, characterized in that: The weight distribution is achieved through the fully connected layer, which specifically includes: dimensionality reduction of the concatenated matrix of the Top-n search results, generating attention weights based on contextual semantics, and adjusting the contribution ratio of different knowledge sources.

Citation Information

Patent Citations

  • LLM capability enhancement method of hybrid enhancement strategy

    CN119578557A

  • Fine adjustment method, system and equipment based on large language model and medium

    CN117290480A

  • Intelligent question answering system based on large language model

    CN119623646A

  • Domain-adaptive retrieval enhancement generation method and system

    CN119669400A

  • Power field knowledge question and answer optimization system based on large model retrieval enhancement generation and instruction supervision fine tuning

    CN119961388A

Cited By

  • Intelligent retrieval method and system for educational resources

    CN121501968A

  • Contract named entity recognition method fusing knowledge base and self-verification mechanism

    CN121525678A