Intelligent question-answering method and system based on semantic vector optimization and dynamic prompting
By designing the semantic vector generation module in a large language model and combining the dynamic prompt mechanism and feature compression module, the semantic vector generation process is optimized, solving the shortcomings of the existing RAG systems in terms of semantic expression, adaptability and matching, and achieving more efficient and accurate retrieval and answer generation.
Patent Information
- Application Number
- CN202510111798.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing search-enhanced generation RAG system has shortcomings in semantic expression ability, scenario adaptability and model training target matching, resulting in low correlation of search results and mismatch of generated answers with user needs.
Design a semantic vector generation module in a large language model, and generate semantic vectors that are more in line with the needs of the search task by deeply understanding the semantic information of user queries. Combining the dynamic prompt mechanism and feature compression module, the semantic vector generation process is optimized to improve the accuracy of search and answer generation.
It improves semantic comprehension ability and retrieval accuracy, ensures that the generated answers match user needs, enhances the adaptability and efficiency of the model, and significantly improves the quality and relevance of the answers in the handling of complex areas.
Smart Images

Figure CN119557407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to an intelligent question-answering method and system based on semantic vector optimization and dynamic prompting. Background Art
[0002] In recent years, large language models (LLMs) have been widely used in the field of natural language processing (NLP), especially in generation tasks (such as text generation, dialogue systems) and knowledge enhancement tasks (such as retrieval-augmented generation (RAG)). RAG combines generative models and retrieval technology, and with the assistance of external knowledge bases, solves the problems of limited knowledge scope and insufficient memory capacity inherent in large language models. However, existing RAG technology still faces many challenges, which limit its widespread deployment and efficient application in practical scenarios.
[0003] In traditional retrieval-augmented generation (RAG) systems, user input is usually generated through the embedding layer of the language model to generate semantic vectors, which are used as query vectors for the retrieval module. However, the embedding layer of the large language model is mainly designed for generation tasks, rather than being optimized specifically for retrieval tasks. Therefore, the vectors generated directly from the embedding layer may have the following problems in expressing user intent and capturing contextual semantic features: Insufficient semantic expression ability: The semantic vectors generated by the embedding layer cannot fully capture the complex semantics in the query, resulting in low relevance of the retrieval results. Limited ability to adapt to scenarios: Existing methods usually use fixed embedding generation methods, which makes it difficult to flexibly adjust the vector generation method according to different task requirements. Model training goal mismatch: The training goal of the large language model focuses on generation capabilities, resulting in the inability of the vectors generated by the embedding layer to fully adapt to the retrieval task requirements.
[0004] In the retrieval-augmented generation (RAG) system, the retrieval module and the generation module are usually separated, and there is a lack of an organic coordination mechanism between the two. This decoupling approach brings the following problems: The retrieval and generation goals are disconnected: The optimization of the retrieval module only considers the retrieval performance of the semantic vector, but fails to consider the specific requirements of the generation module for the retrieval results, which may lead to inconsistency between the retrieved information and the answer requirements of the generation model. Insufficient context fusion: When the retrieved document information is combined with the generation module, the quality of the generated answer may be affected due to semantic incoherence or insufficient context information.
[0005] The semantic vector generation module and generation module in the current retrieval-augmented generation (RAG) method are usually optimized independently and fail to form a closed-loop linkage training, which makes it difficult to improve the overall retrieval and generation performance. The embedded generation process lacks the ability to adapt to the context of the generation task, resulting in the generated answers not matching the actual needs of users.
[0006] Different layers of a large language model encode semantic information of different granularities. Shallow semantics encodes the basic lexical and grammatical structure of the input sentence. Deep semantics captures the contextual semantics of the input sentence and higher-order semantic abstractions. In current methods, retrieval tasks often rely only on shallow embeddings, while generation tasks use more deep semantics. This semantic separation method fails to fully utilize the multi-level semantic information of a large language model. Summary of the invention
[0007] In order to solve the shortcomings of the prior art, the present invention provides an intelligent question-answering method and system based on semantic vector optimization and dynamic prompting; the present invention designs a semantic vector generation module in a large language model to generate a semantic vector that better meets the requirements of the retrieval task by deeply understanding the semantic information of the user's query. This module can not only improve the quality of the embedded vector, but also ensure that the generated vector works effectively with the subsequent generation module, thereby improving the accuracy and relevance of the final answer.
[0008] On the one hand, an intelligent question-answering method based on semantic vector optimization and dynamic prompts is provided, including:
[0009] Get the target question;
[0010] Input the target question into the trained question-answering model to obtain the answer corresponding to the target question; wherein the trained question-answering model includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module;
[0011] The training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module into the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model.
[0012] On the other hand, an intelligent question-answering system based on semantic vector optimization and dynamic prompts is provided, including:
[0013] An acquisition module is configured to: acquire a target problem;
[0014] A question-answering module, which is configured to: input a target question into a trained question-answering model to obtain an answer corresponding to the target question; wherein the trained question-answering model includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer, and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module;
[0015] The training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module into the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model.
[0016] The above technical solution has the following advantages or beneficial effects:
[0017] The present invention adds a semantic vector generation layer module to the large language model to achieve the collaborative work of the semantic vector and the generation module and the dynamic prompt mechanism, thereby solving many shortcomings of traditional methods in processing complex field problems.
[0018] The present invention designs a trainable semantic vector generation module in the embedding generation module of the large language model, which can accurately capture the deep semantic information of the user query and highly match the retrieval task requirements, thereby improving the retrieval accuracy of related documents. This module can not only improve the semantic understanding ability, but also perform adaptive training on various fields (such as law, medical, etc.) to enhance the adaptability of the model. Compared with the traditional fixed embedding generation method, this module can be dynamically adjusted according to data features and context in practical applications, thereby improving the accuracy of retrieval and answer generation. In addition, combined with the design of the dynamic hierarchy selection module and the feature compression module, the semantic vector generation process is further optimized. The dynamic hierarchy selection module enables the model to select the most appropriate hierarchy according to the task requirements by flexibly adjusting the contribution of each layer, thereby extracting the most relevant information in complex queries. The feature compression module ensures that the generated semantic vector reduces the computational overhead while maintaining the richness of information through effective dimensionality reduction processing and nonlinear activation functions, further improving the efficiency and application performance of the model.
[0019] The present invention adopts the collaborative work of semantic vectors and generation modules, breaking the separation of the embedding vector and the generation stage in the traditional architecture, avoiding the disconnection problem between the retrieval stage and the generation stage, so that the final generated answer can better integrate the user's query intention and the retrieved document content, ensuring the contextual coherence and semantic consistency of the answer, thereby providing more accurate contextual information for the final answer generation. Through this collaborative work, the system can handle complex queries and generate high-quality answers that meet the needs of the field. Especially in professional field problems, it can generate high-quality answers that meet the field specifications and have logical consistency, solving the problem of excessive generalization and insufficient field adaptability of traditional generation model answers.
[0020] The present invention adopts a dynamic prompt mechanism. During reasoning, a large language model dynamically generates prompts based on user queries and retrieval content to better adapt to different application scenarios and ensure that the generated prompts are closer to user needs. The dynamic prompt mechanism is particularly effective in multi-domain and multi-task scenarios, ensuring that the generated answers are more context-relevant and domain-specific, which can greatly improve the model's adaptability. The present invention can quickly adapt to different knowledge structures and user needs and generate more targeted answers. This capability not only expands the application scope of the system, but also greatly improves user experience and satisfaction.
[0021] In summary, this solution is expected to achieve breakthrough improvements in retrieval accuracy, answer quality, and domain adaptability by optimizing the generation quality of semantic vectors, improving the synergy between retrieval and generation, and introducing a dynamic prompt mechanism. Compared with the prior art, the present invention demonstrates significant advantages in handling complex domain problems, and can provide users with more accurate, efficient, and domain-specific intelligent question-and-answer services. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0023] Figure 1 is a flow chart of the method of embodiment 1;
[0024] Figure 2 Schematic diagram of the internal structure of the question-answering model of Example 1. DETAILED DESCRIPTION
[0025] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0026] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0027] In this embodiment, all data is obtained in compliance with laws and regulations and based on the user's consent, and is used legally.
[0028] Embodiment 1
[0029] This embodiment provides an intelligent question-answering method based on semantic vector optimization and dynamic prompting;
[0030] like Figure 1 As shown, the intelligent question answering method based on semantic vector optimization and dynamic prompts includes:
[0031] S101: Get the target question;
[0032] S102: Input the target question into the trained question-answering model to obtain the answer corresponding to the target question; wherein the trained question-answering model, such as Figure 2 As shown, it includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module; the training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module into the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model.
[0033] Furthermore, the pre-training of the semantic vector generation module includes:
[0034] Constructing a training set; the training set includes: known questions, correct answers corresponding to the known questions, and incorrect answers corresponding to the known questions;
[0035] The training set is input into the semantic vector generation module, and the semantic vector generation module is trained. When the contrastive loss function (Contrastive Loss) value no longer decreases, the training is stopped to obtain a pre-trained semantic vector generation module.
[0036] Furthermore, the contrastive learning loss function (Contrastive Loss) is expressed as:
[0037]
[0038] in, Represents the contrastive loss function, which calculates the average loss of the entire training batch; Represents the number of samples in the current training batch. The loss calculation in the formula is to sum and average all samples in a batch. Represents the index of a specific sample in the training batch, and its value range is . represents the distance metric between vectors (cosine similarity), represents the label value, indicating the semantic similarity ( Indicates similarity, indicates dissimilarity). represents a hyperparameter that defines the minimum distance between positive and negative samples. Maximize: the semantic similarity between user questions and related content, that is, minimize . Minimize: The semantic similarity between user questions and irrelevant content, that is, maximize .
[0039] Furthermore, the total loss function of the question-answering model includes:
[0040] A multi-task learning framework is adopted to integrate the goals of semantic vector generation and answer generation, and the following joint loss function is designed:
[0041]
[0042]
[0043] Among them, the retrieval loss and generate loss By weighted combination, using hyperparameters and To balance the contribution of both to the final loss function, Controls the weight of the retrieval loss term, Controls the weight of the generated loss term, represents the number of time steps to generate the answer (i.e. the length of the answer), is generated words, It is the retrieved relevant documents, or the input information combined with the dynamic prompt generation mechanism. The generative model generates words given historical words and document information. The probability of retrieval loss Based on the embedding vector and similarity retrieval task, contrastive loss is used to optimize the quality of vector generation and generate loss Calculate the cross entropy loss for the answer of the generation module to ensure that the generated answer meets the requirements of the context, hyperparameters and Dynamically adjust according to task requirements.
[0044] Furthermore, the construction of the semantic vector generation module includes:
[0045] (1-1) The backbone network of the Alibaba Qwen2 model consists of n decoder layers connected in series.
[0046] (1-2) Treat each Decoder layer as a block layer, and get n block layers;
[0047] (1-3) Set the initial weight for each of the n block layers ;
[0048] (1-4) According to the initial weight of each block layer , calculate the probability distribution of each block layer ;
[0049] (1-5) The probability distribution of each block layer Compare with the set threshold. If it is lower than the set threshold, the block layer is removed; if it is higher than the set threshold, the block layer is retained.
[0050] (1-6) All the remaining block layers are connected in series according to the order in the Alibaba Qwen2 model to obtain the semantic vector generation module;
[0051] (1-7) Perform the following operations on each block layer of the semantic vector generation module: Operation to obtain the probability distribution of each block layer of the semantic vector generation module ;
[0052] (1-8) Probability distribution of each block layer based on the semantic vector generation module , and the feature vector output by each block layer of the semantic vector generation module, generate the final embedding representation output by the semantic vector generation module.
[0053] It should be understood that the Ali Qianwen Qwen2 model includes: an input layer, a backbone network, a linear layer, an activation function layer, and an output layer connected in sequence;
[0054] The backbone network includes: an Embedding layer, a first hidden layer, a plurality of serially connected Decoder layers, and a first normalization layer.
[0055] The Decoder layer includes: a second hidden layer, a second normalization layer, an attention mechanism layer, a third hidden layer, a first adder, a fourth hidden layer, a third normalization layer and a multi-layer perceptron layer connected in sequence; the output end of the second hidden layer is residually connected to the input end of the first adder; the output end of the fourth hidden layer is connected to the input end of the second adder; the output end of the second adder is connected to the input end of the fifth hidden layer; the output end of the multi-layer perceptron is connected to the input end of the second adder; the input end of the second hidden layer is the input end of the Decoder layer; the output end of the fifth hidden layer is the output end of the Decoder layer.
[0056] Furthermore, (1-3) sets the initial weight for each of the n block layers. ,include:
[0057] Set an initial weight vector for each Block layer to be screened ,in, Indicates the number of layers involved in the calculation in the model (i.e. the total number of blocks), Indicates The initial weight of layer Block indicates its initial importance in generating the embedding representation.
[0058] Furthermore, (1-4) is based on the initial weight of each block layer , calculate the probability distribution of each block layer ,include:
[0059] The initial weights of each layer conduct Operation, converting it into a probability distribution The specific formula is:
[0060]
[0061] in, Indicates The normalized weight of the block layer indicates the contribution of the block layer to the final embedding representation and satisfies The sum of is 1, is an exponential function. The introduction of exponential function can amplify larger , so that More sensitive to larger weights, is the temperature parameter.
[0062] It should be understood that in order to increase the flexibility of the training process, the present invention uses a temperature scaling mechanism to dynamically adjust the temperature parameters. control The smoothness of the output can flexibly balance the feature contribution of global exploration and local focus. Set to a higher value (such as ),at this time The output is relatively smooth, all weights The distribution is close to uniform, which enables the model to fully explore the potential contributions of different layers and allows the model to explore the role of multiple layers of features. According to the preset index The descent strategy gradually decreases ( is the temperature, is the rate of decrease), when the temperature approaches a lower value (such as T=0.5), The output tends to be sharp, and the weight distribution is concentrated in a few layers, thereby strengthening the contribution of these layers to the embedded representation and intensively utilizing the feature outputs of the key layers.
[0063] Furthermore, the above (1-7) performs the following operations on each block layer of the semantic vector generation module: Operation to obtain the probability distribution of each block layer of the semantic vector generation module ; The specific calculation process is consistent with (2-4).
[0064] Furthermore, the probability distribution of each block layer of the semantic vector generation module (1-8) is , and the feature vector output by each block layer of the semantic vector generation module, generate the final embedding representation output by the semantic vector generation module, including:
[0065] Probability distribution Output feature vector for weighted block layer , generates the final embedding representation output by the semantic vector generation module , the formula is:
[0066]
[0067] Indicates The output feature vector of layer Block, Represents the final generated semantic embedding, which is the weighted sum of the outputs of all layers.
[0068] The semantic vector generation module is implemented by several block layers of the Qwen2 large model that are dynamically selected during training. Here, block refers to the core computing unit in the Qwen2 model. Each block combines the Multi-Head Self-Attention mechanism (MHSA) and the Feedforward Neural Network (FFN). These blocks are the basic components of the Qwen2 model, responsible for extracting and processing the semantic information in the user query layer by layer to generate rich contextual representations. This module deeply understands the user's query, extracts accurate semantic information, and converts it into high-quality semantic vectors. First, the user's query is converted into a low-dimensional dense vector representation by the Embedding Layer. The Embedding Layer maps each query word to a fixed-dimensional word vector through a table lookup operation. These initially generated vectors capture the basic semantic information of each word in the query, but cannot fully express the contextual meaning at the sentence level.
[0069] The dense vector is fed into the block layer of the semantic vector generation module for processing. In these blocks, the multi-head self-attention mechanism calculates the relationship between each word and other words in the sentence to form a context-aware semantic representation; the forward propagation network further performs nonlinear transformations on these context representations to enhance the expressiveness of the features. In this process, features between layers are gradually accumulated to form a high-dimensional vector representation containing deep semantic information. In order to generate deep semantic vectors more efficiently, the semantic vector generation module integrates a dynamic hierarchical selection mechanism. With the support of this mechanism, the model performs a weighted combination of the outputs of all blocks.
[0070] During the training phase, the dynamic level selection module optimizes these weight parameters through back propagation, so that the model can dynamically adjust the feature contribution of each layer according to the characteristics of the input query, thereby ensuring that the generated semantic vector is more adaptable to user queries. In order to optimize the matching degree between the generated semantic vector and the relevant documents, the semantic vector generation module introduces contrastive learning loss during training. Specifically, for each query vector, we maximize its similarity with the relevant document vector and minimize its similarity with the irrelevant document vector. This process can significantly improve the discrimination ability and retrieval performance of the semantic vector.
[0071] Based on a workflow that combines domain dataset training, dynamic level selection, and contrastive loss optimization algorithm, the semantic vector generation module is able to provide accurate and efficient semantic representation for the entire question-answering system, thereby providing more precise and relevant results in complex tasks.
[0072] Furthermore, the embedding layer is implemented by an embedding layer.
[0073] It should be understood that the embedding layer refers to: the embedding layer is the first layer of data processing in the Alibaba Qwen2 series large model, which converts discrete inputs (such as vocabulary) into continuous high-dimensional vectors, and continuously optimizes the representation of these vectors according to contextual relationships during model training. The embedding layer can not only capture semantic information, but also effectively reduce the dimension of data, reduce computational complexity, and provide powerful expressiveness in multiple natural language processing tasks through context-aware learning mechanisms.
[0074] Furthermore, the feature compression module includes: a multi-layer perceptron, an activation function layer and an L2 normalization layer connected in sequence; the multi-layer perceptron is used to perform dimensionality reduction processing on the semantic vector generated by the semantic vector generation module, the activation function layer filters out eigenvalues less than zero, and the L2 normalization layer normalizes the vector.
[0075] It should be understood that the feature compression module performs dimensionality reduction processing on the semantic vectors generated from the semantic vector generation module through a multi-layer perceptron (MLP), thereby generating a more compact and efficient embedding representation and adapting to the fixed embedding dimension requirements of the subsequent RAG index. During the forward propagation of the model, the feature compression module receives high-dimensional vectors generated from the semantic vector generation module, which contain rich contextual information and query semantic features. However, since the dimensions of these hidden states are usually high, directly using these high-dimensional features may lead to low computational efficiency and increase storage overhead. Therefore, the feature compression module maps these high-dimensional features to a lower dimension through an MLP. In order to increase nonlinear characteristics and expressive power, the compressed features are also processed by the ReLU activation function in turn to filter out eigenvalues less than zero, thereby improving the sparsity and recognition ability of the feature representation. Then the feature compression module performs L2 normalization on the activated vector, and the generated embedding vector has a consistent scale, which is convenient for the subsequent RAG module to perform vector matching.
[0076] Furthermore, the RAG retrieval module is implemented by retrieval enhancement generation (RAG), which specifically includes: searching for relevant documents with the highest similarity to the query vector from the existing knowledge base.
[0077] It should be understood that the RAG retrieval module includes: The RAG (Retrieval-Augmented Generation) retrieval module is responsible for retrieving relevant documents based on the query input by the user and providing valuable information support to the answer generation module.
[0078] In the early preparation stage, the system first preprocesses the documents in the knowledge base. This includes slicing the original documents to ensure that the length of each knowledge unit is suitable for the input requirements of the model. The document is segmented using a fixed length or semantic boundary strategy to ensure that each slice contains complete semantic information.
[0079] The segmented documents are passed to the embedding model (bge-large-zh-v1.5) for semantic embedding. The embedding model converts each document slice into a high-dimensional semantic vector to capture its deep semantic features. These vectorized representations are the core of knowledge base construction and are used to support subsequent retrieval tasks. After vectorization, the system uses the Faiss tool to establish efficient indexes for these semantic vectors to form a vector knowledge base. This index library can quickly support large-scale semantic similarity matching and ensure retrieval efficiency.
[0080] When a user sends a query to the system, the RAG module starts the online retrieval process. The user query is first input into the semantic vector generation module to generate a high-quality query semantic vector. This vector not only reflects the literal meaning of the user query, but also deeply explores its potential semantic needs.
[0081] Subsequently, the generated query vector is passed to the retriever, which performs similarity matching in the vector knowledge base based on the Faiss algorithm. Specifically, the retriever quickly selects the most relevant documents to the query by calculating the similarity between the query vector and the document vector (using cosine similarity). During the matching process, the Top-K strategy is adopted to sort the documents from high to low according to the cosine similarity score, and then select the top K documents as the final retrieval results. This ensures that the selected documents are not only highly relevant, but also have information diversity.
[0082] Once relevant documents are retrieved, the RAG module passes them along with the query vector to the dynamic prompt generation module and the answer generation module. These documents provide rich background information for prompt generation and answer generation, ensuring that the system can generate accurate and contextual answers. By combining the query and retrieved document content, the RAG module significantly improves the system's performance in complex tasks and avoids generating empty or general content.
[0083] Furthermore, the dynamic prompt generation module is implemented by an LLM model; the input values of the LLM model are the vector to be queried and the retrieved related documents, and the output value of the LLM model is the prompt word for generating the vector to be queried.
[0084] It should be understood that the dynamic prompt generation module calls the large model (LLM) to dynamically generate personalized and efficient prompts (Prompt) through user queries and retrieved related documents to guide the dynamic prompt generation module to provide more accurate and user-friendly answers. After the user submits a query request, the system first understands the semantics of the query through the semantic vector generation module and generates a corresponding high-quality semantic vector. The vector is then used to find documents that are highly relevant to the query semantics through the retrieval enhancement generation (RAG) module. These retrieved documents are sorted, screened, and deduplicated to ensure that the content of subsequent processing is highly relevant and consistent.
[0085] After receiving the user's query and the retrieved documents, the dynamic prompt generation module will conduct an in-depth analysis of this information, including the document category, content characteristics, and the semantic relevance to the user's question. For example, when analyzing the document category, the module can dynamically adjust the prompt generation style based on the document's source or subject (such as a technical manual, legal text, or news report); when analyzing the content characteristics, the module will extract key sentences, core topics, or statistical features in the document to enhance the background information of the prompt. Through this analysis process, the system can effectively extract key information that can support answer generation.
[0086] Based on the above analysis results, the dynamic prompt generation module starts to build prompts. The generation of prompts starts with basic templates, which contain common instruction structures and question background frameworks. Then, the module dynamically injects important semantic information in the user query and key content in the retrieved document, and generates the final prompt by fusing these elements. Specifically, prompt generation includes the following steps: first, providing background information for the prompt based on the context of the user query; second, extracting supporting content from the document and forming the main part of the prompt by insertion or reorganization; third, strengthening the generation goal through instructions, such as clarifying the format, language style or depth requirements of the answer.
[0087] When calling LLM to generate prompts, the system uses the information obtained from the analysis to dynamically build the model input. For example, the input format can be a sequence containing user questions, document summaries, and generation goals, with the following specific format:
[0088] User query: {user's question};
[0089] Retrieve documents: {document summary or key sentences};
[0090] Generation goal: {answer style, answer format, or depth requirement}.
[0091] Through this input format, LLM can generate prompts with high semantic relevance. Finally, these dynamic prompts are passed to the generation module to guide the model to output answers that meet user needs. The dynamic prompt generation module can adjust the prompt content in real time according to different tasks and query requirements, thereby significantly improving the overall performance of the question-answering system in terms of context adaptability, answer accuracy and generation efficiency.
[0092] Furthermore, the answer generation module is implemented by n decoder layers connected in series in the backbone network of the Alibaba Qwen2 model.
[0093] The answer generation module is responsible for generating answers based on the user's query and the retrieved relevant documents.
[0094] Generate high-quality, accurate and contextually coherent answers based on the user's query and the retrieved relevant documents. The answer generation module is based on the basic block architecture of the Qwen2 model, based on dynamic prompt guidance, through deep information fusion and attention mechanism, to ensure that the generated answers can meet the needs of complex tasks.
[0095] The answer generation module receives the optimization hints generated by the dynamic hint generation module and the retrieved related documents as input. The optimization hints include the semantic information of the user query, the background content of the retrieved documents, and the target instructions of the answer.
[0096] The core of the answer generation module is the multi-head self-attention mechanism and feedforward network of the decoder layer in the Qwen2 model. Through multi-layer stacked blocks, the model can perform deep semantic fusion and reasoning on input prompts and document content.
[0097] In terms of context processing, the model captures and comprehensively analyzes remote dependencies through the semantic modeling capabilities of multiple layers of blocks. In this process, the generated goals in the prompts and the content of the retrieved documents are gradually integrated to form a coherent semantic representation. Finally, the output layer decodes the intermediate representation into a specific answer text.
[0098] In addition, the answer generation module dynamically adjusts the generation strategy according to the question type. For open-ended questions, the module tends to generate rich and comprehensive answers; for questions in specific fields (such as law or medicine), the module focuses on providing accurate answers that are in line with the field background. Through dynamic prompt guidance, deep information fusion and context analysis, the generated answers can accurately respond to user needs while maintaining the logic and fluency of language expression. Finally, the answer generation module feeds back the generated answers to the user, completing the entire question-and-answer process. This module provides efficient and reliable solutions to complex problems by optimizing the generation process and model structure.
[0099] Furthermore, the working process of the trained question-answering model includes:
[0100] The feature compression module performs dimensionality reduction processing on the semantic vector generated by the dynamic hierarchical selection module and compresses it into the target vector space, reducing the computational complexity while retaining the core information.
[0101]
[0102] Represents a feature compression function, which is used to convert high-dimensional semantic features Compressed into low-dimensional semantic features of the target vector space , It is the weighted hidden state output by the dynamic level selection module, which is an important feature selected from several layers of the large model through weight distribution. W represents the weight matrix, which is a learnable parameter used for linear transformation. Represents the bias vector, which is also a learnable parameter used to increase the expressiveness of the model. Represents the activation function.
[0103] The semantic vector generation module performs semantic understanding of the target problem and generates the target semantic vector through the dynamic hierarchical selection module and the feature compression module.
[0104]
[0105] The RAG retrieval module calculates the similarity between the target semantic vector and the documents or fragments in the vector knowledge base, and retrieves the content with similarity higher than the threshold.
[0106]
[0107] is the similarity function, which represents the cosine similarity between two vectors. Represents a vector in the knowledge base. is the similarity score of two vectors.
[0108] The dynamic prompt generation module generates context-enhanced prompt content based on the retrieval results, target questions and related fields, providing background information for the answer generation module.
[0109]
[0110] is the collection of retrieved documents, The domain context representing the target problem, Indicates prompt word generation.
[0111] The answer generation module generates the final answer based on the target semantic vector, retrieval results and context-enhanced prompt content.
[0112]
[0113] is the output of the dynamic prompt generation module, is the collection of retrieved documents, is the target semantic vector.
[0114] It should be understood that the semantic vector generation module performs deep semantic understanding of the user's input questions. The semantic vector generation module combines multiple layers of semantic information to generate high-quality semantic vectors for input questions. These vectors can not only fully capture the core intention of the user's questions, but also show excellent retrieval capabilities. The semantic vector generation module is closely integrated with the embedding layer and hidden layer of the basic model through a dynamic hierarchical selection mechanism, dynamically selects the most semantically representative hidden state features, and learns the importance of semantic features of each layer through a weight distribution strategy, thereby ensuring efficient fusion of semantic information. The implementation of the semantic vector generation module relies on a dynamic weight distribution strategy, using a Softmax function with temperature control to dynamically adjust the layer weight distribution according to feedback during the training process. In addition, the semantic vector generated by the semantic vector generation module is also processed by a feature compression module for dimensionality reduction, further improving the retrieval performance and computational efficiency of the vector. The generated semantic vector is designed as a multifunctional vector, which is not only suitable for text retrieval, but also directly enhances the semantic vector generation module's ability to understand user needs, laying the foundation for the execution of subsequent modules.
[0115] It should be understood that the feature compression module is responsible for mapping the high-dimensional semantic vector generated by the layer selector to a more compact low-dimensional semantic space to optimize the embedding quality and reduce computational overhead. This module combines the multi-layer perceptron (MLP) structure, nonlinear activation function and L2 normalization processing to dynamically compress the semantic vector. Through this design, the feature compression module significantly improves the computational efficiency and retrieval adaptability of the vector while maintaining the integrity of the semantic information. In addition, the semantic vector after feature compression not only improves the similarity matching accuracy with the vector in the knowledge base, but also provides more targeted semantic input for the subsequent generation module, thereby enhancing the overall performance and robustness of the system.
[0116] It should be understood that the RAG retrieval module is responsible for using semantic vectors to retrieve documents or fragments that are highly relevant to the question in the knowledge base. This module uses the Faiss algorithm, combined with dynamically optimized embedding vectors and vectors in the vector knowledge base for similarity matching, to accurately retrieve specific field questions. By retrieving relevant documents, this module provides the necessary content support for the dynamic prompt generation module, and also improves the answer quality and context adaptability of the generation module.
[0117] It should be understood that the dynamic prompt generation module is designed to improve the accuracy and pertinence of answers. After obtaining the relevant documents retrieved by the RAG module, it combines user questions and domain requirements to generate a dynamically adjusted prompt. Through the prompt dynamically constructed by LLM, this module strengthens the model's ability to understand specific scenarios and questions, making the final generated answer more in line with the actual needs of users and having good domain adaptability.
[0118] It should be understood that the answer generation module is responsible for generating the final answer. It generates high-quality answers that meet user needs by combining the understanding results of the semantic vector generation module, the search content of the RAG module, and the context-enhanced prompts of the dynamic prompt module. This module further optimizes the fluency and accuracy of the generation, ensuring that the answer is not only in line with the field background, but also has clear and professional expression.
[0119] The semantic vector generation module, feature compression module, RAG retrieval module, dynamic prompt generation module, and answer generation module, these five modules achieve efficient operation of the system through close collaboration. The semantic vector generation module performs deep semantic analysis on the questions input by the user to generate high-quality semantic vectors. These vectors not only reflect the core needs of the user's questions, but also provide strong semantic support for subsequent steps. The generated semantic vectors are compressed by the feature compression module and then passed to the RAG retrieval module. The RAG retrieval module quickly locates the text or information fragments related to the question in the knowledge base to ensure that the system obtains content that can directly serve the user's needs. Then, the retrieved content is sent to the dynamic prompt generation module together with the user's question. By comprehensively analyzing the semantics of the user's question and the retrieved text category, the module dynamically generates a prompt that adapts to the current scenario. This process strengthens the system's understanding depth and generation ability of user questions. Finally, the answer generation module combines the context guidance of the dynamic prompt, the retrieved information and the deep understanding results of the semantic vector generation module to generate high-quality answers, ensuring that the answers meet the actual needs of the user and are accurate and fluent. This multi-module collaboration method enables the system to achieve closed-loop optimization from problem understanding, relevant content extraction to high-quality answer generation, demonstrating a high degree of modular division of labor and collaborative efficiency.
[0120] We selected an existing large language model (based on the Qwen2 model) and redesigned the model structure by introducing a dynamic hierarchical selection mechanism and a feature compression module. Specifically, a semantic vector generation module is added after the input embedding layer of the model, so that the model can further complete deep semantic analysis after generating the initial embedding representation and generate more expressive optimized semantic vectors. The new model structure selects the appropriate model level and assigns weights through a dynamic hierarchical selection mechanism, and optimizes the vectors in combination with a feature compression module, which effectively enhances the model's ability to understand the semantics of user input, especially in complex domain problems. This optimized semantic vector is more suitable for subsequent relevance retrieval and answer generation tasks.
[0121] The user's question is first input into the embedding layer of the model to generate an initial embedding representation. The initial embedding representation is then sent to the semantic vector generation module, where the dynamic level selection mechanism adaptively selects different levels of the model through learning weights to obtain the optimal semantic features; then, the feature compression module performs dimensionality reduction on the generated semantic vector to ensure that the embedding vector is efficient and adaptable. These optimized semantic vectors provide high-quality semantic information support for subsequent relevance retrieval and answer generation tasks.
[0122] The optimized semantic vectors are input into the RAG retrieval module, which finds documents or content blocks related to the user's question in the knowledge base through similarity matching. The knowledge base can be domain-specific text data, structured information, or mixed data. The retrieved documents are filtered and sorted, and the most relevant content is selected for subsequent generation tasks to ensure that the generation process can be based on high-quality supporting information.
[0123] According to the retrieved relevant documents, the prompt content is dynamically generated. The prompt generation process combines user questions and search content classification to adapt to the needs of different scenarios. The dynamic prompt mechanism ensures that the prompt content can guide the large model to focus on the semantic core and background information of the user's questions, and generate more accurate and targeted answers in combination with the search content. The user question, the dynamically generated prompt, and the retrieved relevant documents are input into the answer generation module. The answer generation module generates high-quality answers in combination with the input content. In this process, the dynamic prompt mechanism and the search enhancement generation work together to ensure that the generated answers are both accurate and in line with the context. The question-answering system architecture based on semantic vector optimization and dynamic prompt mechanism of the present invention makes innovative improvements on the basis of the traditional RAG framework by designing a semantic vector generation module and citing a dynamic prompt mechanism. Under the designed question-answering system framework with a semantic vector generation module, a dynamic hierarchical selection mechanism and a feature compression module are designed, which not only optimizes the generation quality of the semantic vector, but also significantly improves the retrieval efficiency and relevance, while the dynamic prompt mechanism further enhances the pertinence and quality of the answer, thereby realizing an efficient question-answering system for complex domain problems. The model level is adaptively selected through a dynamic level selection mechanism, and the feature compression module is used to improve vector expression capabilities.
[0124] The dynamic level selection mechanism can adaptively select the optimal model level combination for generating semantic vectors based on the characteristics of the input question by adding a learnable weight allocation module between specific levels of the model. Specifically, the system will first perform a preliminary feature analysis on the input question, extract key information such as its complexity and context dependency, and dynamically allocate weights in the model's levels based on this information. The learning of weights is completed through the training phase, and the system can gradually optimize the weight allocation strategy to achieve fast and efficient dynamic level selection in the inference phase. In this way, the semantic vector generation module can more accurately capture the semantic features of user questions and improve representation capabilities.
[0125] The feature compression module improves the expressiveness and computational efficiency of semantic vectors by reducing their dimensionality. The module first analyzes the semantic vector of the input question in a high-dimensional space, and retains the key dimensions that best represent the semantic information through a feature selection algorithm (for example, based on principal component analysis PCA or deep adaptive feature extraction). At the same time, the feature compression module further optimizes the representation capability of the remaining features using nonlinear transformation methods, so that the compressed semantic vector can reduce computational overhead while retaining enough information in a low-dimensional space. This compression method is particularly suitable for the needs of efficient retrieval in the RAG framework, and can reduce resource usage while improving retrieval performance.
[0126] Combined with the RAG framework, we can fully utilize the retrieval advantages of the external knowledge base, and ensure the adaptability and contextual relevance of the generated answers through a dynamic prompt mechanism, thereby achieving efficient processing of complex question-answering tasks.
[0127] Generate a domain-specific semantic vector training set to effectively train the semantic vector generation module so that it can generate high-quality semantic vectors for distinguishing relevant and irrelevant content.
[0128] For specific fields (such as law), we first collect relevant data sets from public resources. We obtain the following data sets, Chinese-laws-pretrain, Chinese_Law, and Chinese-law-and-regulation, through the Hugging Face platform. In order to ensure that the data sets cover a wider range of scenarios and more complex types of questions, we generate data through chatgpt. The specific steps are as follows: Generate a fixed-format data set by modifying the prompt words of chatgpt.
[0129] Example of prompt words: Generate one relevant content and one irrelevant content based on the following questions:
[0130] question: {query};
[0131] Output format: {"query":"question", "positive":"relevant content", "negative":"irrelevant content"};
[0132] Sample input:
[0133] Generate one relevant piece of content and one irrelevant piece of content based on the following questions:
[0134] Question: What are the provisions of the Company Law regarding the forms of shareholder capital contribution?
[0135] Output format: {"query":"question","positive":"relevant content","negative":"irrelevant content"}; The collected data sets often contain duplicate, incomplete or irrelevant content to the target field, so the following steps are required for data cleaning:
[0136] Deduplication and completion: remove duplicate records and complete missing content.
[0137] Format standardization: Unify the data format to ensure that each record contains complete necessary fields, such as questions (query), relevant content (positive document), and irrelevant content (negative document).
[0138] Domain consistency: Eliminate records that do not conform to the target domain to ensure that the data strictly belongs to the selected domain.
[0139] After data cleaning, the data needs to be adjusted to a fixed format suitable for the training of the semantic vector generation module, including query (question), positive (relevant content), and negative (irrelevant content).
[0140] The constructed dataset will be used to train the semantic vector generation module. By comparing positive and negative content and combining the dynamic hierarchical selection module with the feature compression module, the semantic vector generation module can generate high-quality semantic vectors, thereby improving the retrieval accuracy and question-answering performance in the retrieval augmented generation (RAG) framework.
[0141] Based on the constructed dataset, the aim is to train the copied semantic vector generation module so that it can generate high-quality semantic vectors. The training goal is to maximize the semantic similarity between user questions (queries) and related content (positive), while minimizing the semantic similarity between user questions (queries) and irrelevant content (negative). To this end, contrastive learning loss is introduced as an optimization target.
[0142] The generated semantic matching dataset is used as the training dataset, where each record contains the following fields:
[0143] query: The question or query entered by the user.
[0144] Positive (relevant content): document content that is highly relevant to the question semantics.
[0145] Negative (irrelevant content): document content that is irrelevant or has low relevance to the question semantics.
[0146] During training, the input consists of three parts: query, positive, and negative. After the three pass through the embedding layer and the semantic vector generation module, three semantic vectors are generated respectively: A semantic vector representing the user's query. Semantic vector representing relevant content (positive). Semantic vector representing irrelevant content (negative).
[0147] Select Contrastive Loss to optimize the model. Its goal is to use the distance metric method to: and The semantic vector distance of and The semantic vector distance is as far as possible.
[0148]
[0149] represents the distance metric between vectors (cosine similarity), is the tag value, represents semantic similarity (1 represents similarity, 0 represents dissimilarity), is a hyperparameter that defines the minimum distance between positive and negative samples. Maximize: the semantic similarity between user questions and related content, that is, minimize . Minimize: The semantic similarity between user questions and irrelevant content, that is, maximize .
[0150] First, load the data and batch it into training and validation sets, and randomly divide the training data for model training. Load the data in batches to ensure that there are positive and negative samples in each batch.
[0151] Then, the semantic vector is generated. Each sample passes through the embedding layer and the semantic vector generation module to generate , , Three semantic vectors. Calculate the loss value L of the current batch according to the loss function (Contrastive Loss). Use the optimizer (Adam) to update the model parameters and gradually reduce the loss. Finally, evaluate the model performance on the validation set and measure the quality of the generated semantic vectors by cosine similarity or Euclidean distance. When the loss value on the validation set converges or reaches the preset number of training rounds, terminate the training.
[0152] The semantic vectors generated by the trained semantic vector generation module are evaluated with the following indicators: Semantic similarity accuracy: correct matching and Semantic Distinction: Successful Distinction and Embedding vector quality: Verify whether the semantic vector generation module generates high-quality semantic vectors through retrieval tasks (RAG framework) to improve the retrieval accuracy and answer quality of the question-answering system.
[0153] The trained semantic vector generation module flexibly combines levels through a dynamic level selection mechanism and generates high-quality semantic vectors in combination with a feature compression module. This module significantly improves the retrieval relevance and accuracy of the RAG framework, and provides a solid foundation for dynamic prompt generation and answer generation in the question-answering system.
[0154] The dynamic prompt mechanism dynamically generates adaptive prompts by combining user queries and retrieved related document categories. The dynamic prompt mechanism aims to guide the model to generate answers that are more in line with domain requirements and context, thereby improving the accuracy and practicality of the Q&A system in specific scenarios. Based on the user query and the retrieved document content, prompts that adapt to domain requirements are generated, providing context-rich guidance information. Dynamic prompts not only enhance the model's understanding of the question context, but also provide targeted answers in complex scenarios.
[0155] The dynamic prompt mechanism generates high-quality prompts by calling the large model (LLM) to comprehensively analyze the user's original question and the retrieved document content, avoiding the limitations of fixed prompts and making the system more flexible to respond to diverse queries.
[0156] The system accepts the user's query and the relevant documents returned by the RAG retriever. The RAG retriever extracts the most relevant content to the user's question from the knowledge base and classifies the documents, such as legal provisions, case analysis, theoretical knowledge, etc.
[0157] The dynamic prompt generation module is implemented by LLM. The input includes: user query (Query), which is the core question of the prompt. Retrieval results, including document content and classification information, are used to generate background information and key content of the prompt. The large language model extracts the subject and classification of the retrieved document (such as regulations and case analysis). Generates background descriptions related to the user query based on the retrieved content. Provides specific generation targets and context restrictions for subsequent generation modules.
[0158] Example of generating a prompt: "Combining the provisions on breach of contract liability in Article 107 of the Contract Law and the calculation methods in actual cases A and B, analyze the calculation of breach of contract liability in contract disputes. Please provide users with specific compensation standards, applicable scenarios, and relevant explanations."
[0159] The dynamic prompt mechanism will dynamically adjust the prompt content according to the classification information of the retrieved document.
[0160] For example, if the search documents are mainly legal provisions, Prompt prefers to cite legal provisions and legal definitions.
[0161] If the retrieved document is a case study, Prompt focuses on practical applications and detailed explanations.
[0162] The search results usually contain multiple documents. The prompt generation module integrates the main information of these documents to ensure that the generated prompt is comprehensive. The prompt content adapts to the user's question context to ensure that the model understands the user's true intention. The dynamically generated prompt is passed to the subsequent LLM layer together with the retrieved documents. LLM generates answers that meet the needs of specific scenarios based on the context and guidance provided by the dynamic prompt.
[0163] By training the prompt generation module, the prompt guidance performance is maximized, so that the answer is highly matched with the user's needs. A dataset containing questions, search content and high-quality prompts is constructed to fine-tune the prompt generation capability of LLM. Dynamic prompts incorporate the background of retrieved documents into the question answering process, so that the model can better meet the needs of specific fields (such as law and medicine). The contextual information in the prompt makes the model's answer more targeted, avoiding answers that are too general or out of touch with reality. The dynamic prompt mechanism supports multi-scenario applications, without the need to design a prompt template for each scenario separately, greatly improving the scalability of the question-answering system.
[0164] A multi-task learning framework is built to integrate the goals of the semantic embedding module and the dynamic prompt generation module (LLM) to achieve collaborative optimization. The optimized embedding vector can not only improve the relevance of the retrieval stage, but also directly enhance the context understanding ability and answer quality of the generation module, thereby improving the performance of the overall question-answering system.
[0165] Embedding vectors can better express the semantic features of user questions and form stronger semantic associations with the retrieval content. Semantic vectors provide high-quality input context for the generation model, thereby reducing bias and redundancy during generation. Through the multi-task learning framework, semantic vector generation is integrated with answer generation, and embedding vectors are optimized so that they are not only suitable for retrieval, but also directly improve the context quality of the generation module. The generation module can make full use of the retrieved information, improve the synergy of the two modules, and improve the completeness and accuracy of the answers.
[0166] The goal of joint training is to make the semantic vector generation module and the answer generation module work together under the same framework to achieve the optimization of the semantic vector and the improvement of the generation performance.
[0167] In the generation task, the output vector of the semantic vector generation module is directly used as an additional input of the generation module to enhance the semantic relevance of the generation process.
[0168]
[0169]
[0170] In the above code, when the model receives the concatenated input content (query + retrieved_docs), it also introduces the high-quality semantic vector (semantic_vector) generated by the semantic vector generation module. This design allows the semantic vector to not only participate in the retrieval task, but also directly affect the decision-making process of the generation model. query: the user's query input, that is, the question to be answered. retrieved_docs: the relevant documents returned by the RAG retrieval module as supplementary context information. tokenizer: used to convert text input into a tensor format that the model can handle. return_tensors="pt" means returning a PyTorch tensor. inputs: indicates that the model input after the user query and the retrieved document are concatenated, which has been converted into a tensor format. embeddings=semantic_vector: additional semantic vector input provided to the generation model. The semantic vector is generated by the semantic vector generation module to provide additional semantic guidance for the generation process. logits: the output score of the model, used to generate the final answer.
[0171] During the joint training process, the loss signal is transmitted to the semantic vector generation module, including the dynamic layer selection mechanism and the feature compression module, through back propagation in the generation task. The dynamic layer selection mechanism dynamically allocates the semantic contributions of different layers by learning the adjustable weights of each layer; the feature compression module optimizes the semantic vector by dimensionality reduction to improve its adaptability and computational efficiency. This joint optimization strategy significantly improves the adaptability and expressiveness of the semantic vector for the generation task.
[0172] The feature compression module is responsible for mapping the high-dimensional semantic vector generated by the layer selector to a more compact low-dimensional space to optimize the embedding quality and reduce the computational overhead. The implementation of this module includes a multi-layer perceptron (MLP), a nonlinear activation function (ReLU), and L2 normalization. The MLP structure projects the high-dimensional semantic vector (such as 1024 dimensions) through a multi-layer perceptron and maps it to a predefined low-dimensional space (such as 768 dimensions). This multi-layer structure can capture more inter-feature relationships and improve nonlinear expression capabilities. The semantic vector transformed by MLP will be activated by the ReLU activation function to introduce nonlinear characteristics, which enhances the flexibility of the model, especially showing stronger fitting ability when processing complex semantics. The ReLU function increases sparsity and improves the training effect of the network by truncating negative values. Finally, the vector after nonlinear activation will be L2 normalized to standardize the modulus of the vector, ensure the balanced influence of different dimensions, and help improve the stability of the system and optimize the computational efficiency. The output of the feature compression module not only retains the core information of the semantic vector, but also significantly reduces the demand for computing resources in subsequent retrieval and generation tasks.
[0173] The dynamic layer selection module, feature compression module and semantic vector generation module are independent and highly collaborative in the system architecture to jointly complete the generation and optimization of semantic vectors. The dynamic layer selection module dynamically selects certain layers of the deep model through the layer selector, and flexibly adjusts the weight distribution of each layer according to task requirements, thereby generating high-quality semantic representation vectors.
[0174] The output semantic vector of the semantic vector generation module is passed to the feature compression module for feature dimension reduction and structure optimization. The feature compression module is based on multi-layer perceptron, nonlinear activation function and L2 normalization, and maps the high-dimensional semantic vector to a low-dimensional embedding space to adapt the RAG index, which reduces storage and computing costs while retaining key semantic features. Through this processing, the generated compact semantic vector can not only reduce the interference of irrelevant features, but also improve the efficiency and performance in embedding matching and retrieval tasks.
[0175] In general, the dynamic level selection module is responsible for dynamically selecting the appropriate block layer from the deep model, the feature compression module performs dimensionality reduction optimization on the generated semantic vector, and the semantic vector generation module seamlessly connects the output of the dynamic level selection module with the input of the feature compression module to ensure that the generation and processing of the semantic vector are highly coordinated. The organic combination of the three not only optimizes the transmission and expression of semantic information, but also significantly enhances the adaptability and efficiency of the question-answering system in complex field problem processing.
[0176] Embodiment 2: This embodiment provides an intelligent question-answering system based on semantic vector optimization and dynamic prompts, including: an acquisition module, which is configured to: acquire a target question;
[0177] A question-answering module, which is configured to: input a target question into a trained question-answering model to obtain an answer corresponding to the target question; wherein the trained question-answering model includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer, and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module;
[0178] The training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module into the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model.
[0179] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent question-answering method based on semantic vector optimization and dynamic prompts, characterized by: include: Get the target question; Input the target question into the trained question-answering model to obtain the answer corresponding to the target question; wherein the trained question-answering model includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module; The training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module to the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model; Wherein, the semantic vector generation module is constructed, including: (1-1) The backbone network of the Alibaba Qwen2 model consists of n decoder layers connected in series. (1-2) Treat each Decoder layer as a block layer, and get n block layers; (1-3) Set the initial weight for each of the n block layers ; (1-4) According to the initial weight of each block layer , calculate the probability distribution of each block layer ; (1-5) The probability distribution of each block layer Compare with the set threshold. If it is lower than the set threshold, the block layer is removed; if it is higher than the set threshold, the block layer is retained. (1-6) All the remaining block layers are connected in series according to the order in the Alibaba Qwen2 model to obtain the semantic vector generation module; (1-7) Perform the following operations on each block layer of the semantic vector generation module: Operation to obtain the probability distribution of each block layer of the semantic vector generation module ; (1-8) Probability distribution of each block layer based on the semantic vector generation module , and the feature vector output by each block layer of the semantic vector generation module, generate the final embedding representation output by the semantic vector generation module.
2. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: The pair The semantic vector generation module is pre-trained, including: Constructing a training set; the training set includes: known questions, correct answers corresponding to the known questions, and incorrect answers corresponding to the known questions; The training set is input into the semantic vector generation module, and the semantic vector generation module is trained. When the contrastive learning loss function value no longer decreases, the training is stopped to obtain a pre-trained semantic vector generation module.
3. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 2, characterized in that: The contrastive learning loss function is expressed as: ; in, represents the contrastive learning loss function, which calculates the average loss of the entire training batch; Represents the number of samples in the current training batch. The loss calculation in the formula is to sum and average all samples in a batch; Represents the index of a specific sample in the training batch, and its value range is ; represents the distance measure between vectors, represents the label value, indicating the semantic similarity, Represents a hyperparameter that defines the minimum distance between positive and negative samples.
4. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: The total loss function includes: A multi-task learning framework is adopted to integrate the goals of semantic vector generation and answer generation, and the following joint loss function is designed: ; ; ; Among them, the retrieval loss and generate loss By weighted combination, using hyperparameters and To balance the contribution of both to the final loss function, Controls the weight of the retrieval loss term, Controls the weight of the generated loss term, represents the number of time steps to generate the answer, is generated words, are the relevant documents retrieved, The generative model generates words given historical words and document information. The probability of retrieval loss Based on the embedding vector and similarity retrieval task, contrast loss is used to optimize the quality of vector generation and generation loss Calculate the cross entropy loss for the answer of the generation module to ensure that the generated answer meets the requirements of the context, hyperparameters and Dynamically adjust according to task requirements.
5. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: According to the initial weight of each block layer , calculate the probability distribution of each block layer ,include: The initial weights of each layer conduct Operation, converting it into a probability distribution The specific formula is: ; in, Indicates The normalized weight of the block layer indicates the contribution of the block layer to the final embedding representation and satisfies The sum of is 1, is an exponential function, is the temperature parameter.
6. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: Probability distribution of each block layer based on the semantic vector generation module , and the feature vector output by each block layer of the semantic vector generation module, generate the final embedding representation output by the semantic vector generation module, including: Probability distribution Output feature vector for weighted block layer , generates the final embedding representation output by the semantic vector generation module , the formula is: ; Indicates The output feature vector of layer Block, Represents the final generated semantic embedding, which is the weighted sum of the outputs of all layers.
7. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: The embedding layer is implemented by an embedding layer; The feature compression module includes: a multi-layer perceptron, an activation function layer and an L2 normalization layer connected in sequence; the multi-layer perceptron is used to perform dimensionality reduction processing on the semantic vector generated by the semantic vector generation module, the activation function layer filters out eigenvalues less than zero, and the L2 normalization layer normalizes the vector.
8. The intelligent question-answering method based on semantic vector optimization and dynamic prompting as claimed in claim 1, characterized in that: The RAG retrieval module is implemented through retrieval enhancement generation, which specifically includes: searching for relevant documents with the highest similarity to the query vector from the existing knowledge base; The dynamic prompt generation module is implemented by the LLM model; the input value of the LLM model is the vector to be queried and the retrieved related documents, and the output value of the LLM model is the prompt word for generating the vector to be queried.
9. An intelligent question-answering system based on semantic vector optimization and dynamic prompts, characterized by: include: An acquisition module is configured to: acquire a target question; A question-answering module, which is configured to: input a target question into a trained question-answering model to obtain an answer corresponding to the target question; wherein the trained question-answering model includes: an input layer, an embedding layer, a semantic vector generation module, a feature compression module, a dynamic prompt generation module, an answer generation module, a linear layer, and an activation function layer connected in sequence; wherein the output end of the feature compression module is also connected to the input end of the RAG retrieval module, and the output end of the RAG retrieval module is connected to the input end of the dynamic prompt generation module; The training process of the trained question-answering model includes: constructing a semantic vector generation module and pre-training the semantic vector generation module; setting the pre-trained semantic vector generation module to the question-answering model, training the question-answering model, and stopping the training when the total loss function value no longer decreases to obtain the trained question-answering model; Wherein, the semantic vector generation module is constructed, including: (1-1) The backbone network of the Alibaba Qwen2 model consists of n decoder layers connected in series. (1-2) Treat each Decoder layer as a block layer, and get n block layers; (1-3) Set the initial weight for each of the n block layers ; (1-4) According to the initial weight of each block layer , calculate the probability distribution of each block layer ; (1-5) The probability distribution of each block layer Compare with the set threshold. If it is lower than the set threshold, the block layer is removed; if it is higher than the set threshold, the block layer is retained. (1-6) All the remaining block layers are connected in series according to the order in the Alibaba Qwen2 model to obtain the semantic vector generation module; (1-7) Perform the following operations on each block layer of the semantic vector generation module: Operation to obtain the probability distribution of each block layer of the semantic vector generation module ; (1-8) Probability distribution of each block layer based on the semantic vector generation module , and the feature vector output by each block layer of the semantic vector generation module, generate the final embedding representation output by the semantic vector generation module.
Citation Information
Patent Citations
Training method for improving accuracy of large language model in answering legal compliance questions
CN118152628A
Medical literature intelligent question answering system and method based on RAG and LLM technologies
CN118364088A