Large model retrieval enhancement generation method and system based on diffusion alignment
By introducing a diffusion alignment-based method in the large-model retrieval enhancement generation technology, the problem of insufficient inference speed and semantic understanding in the prior art is solved, and more efficient information recall and more accurate semantic understanding are achieved.
Patent Information
- Application Number
- CN202411819577.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-11
Smart Images

Figure CN119988596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model retrieval, and in particular to a large model retrieval enhancement generation method and system based on diffusion alignment. Background Art
[0002] Retrieval-augmented Generation (RAG) leverages the contextual learning capabilities of large language models to enable large language models to acquire the latest knowledge without training, avoiding the defect of acquiring the latest knowledge by retraining the model to update parameters, and greatly reducing the consumption of computing resources without reducing model performance. RAG is usually implemented by using a vector model to encode the latest knowledge, and then saving the encoded information into a vector database; during the reasoning process, the vector model is used to encode the user's question, and then the encoded vector is calculated with the vector in the vector database. The method usually used is Euclidean distance or cosine similarity. Then the text information with high similarity is returned as part of the model input, alleviating the difficulty of not being able to answer new questions due to fixed model parameters.
[0003] One of the key points of RAG technology is to accurately and quickly recall relevant text information based on the user's questions. Since the large language model needs to go through steps such as retrieval and similarity calculation before it is generated, the reasoning speed is greatly reduced compared to general methods; in addition, since the embedding model of the vector model is different from the embedding model of the large language model, the information that the vector model considers relevant may not be relevant in the eyes of the large language model, that is, the retrieval model and the large language model are not aligned, and there are differences in semantic understanding; even if the relevant information can be accurately recalled, it is impossible to eliminate the human errors or data errors caused by building the database, which will seriously affect the output results of the large language model (hallucination, toxicity, instability, etc.). How to accurately recall relevant semantic information while ensuring the reasoning speed is still a problem that needs to be solved. Summary of the invention
[0004] In order to solve the above technical problems existing in the prior art, the present invention proposes a large model retrieval enhancement generation method and system based on diffusion alignment to solve the above technical problems.
[0005] According to a first aspect of the present invention, a large model retrieval enhancement generation method based on diffusion alignment is proposed, comprising:
[0006] S1: Obtain the user's question text information, encode it using a vector model to obtain a semantic vector, calculate the cosine similarity between the semantic vector and the vectors obtained from the vector database in sequence, and extract the text information and vector information whose similarity meets the threshold;
[0007] S2: Extract entities from text information, define entity information categories and make them consistent with the definitions in the graph database, obtain entity sequences, index the extracted entity information in the graph database, obtain additional semantic information and corresponding encoding vectors based on entities and entity relationships, use the encoding vectors as known information to obtain entity-enhanced graph database information, and then merge the graph database information with the information in the vector database to obtain the final background information;
[0008] S3: diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model, introduce preset parameters and calculate the diffusion offset vector;
[0009] S4: continue to input the diffusion offset vector and the updated vector into the semantic diffusion structure, repeat S3 to obtain the information diffusion offset vector, subtract the diffusion offset vector from the information diffusion offset vector bit by bit to obtain the semantic vector that eliminates the pseudo-semantic noise, and repeat until the number of steps reaches the preset value to obtain the final diffusion offset vector;
[0010] S5: Encode the user's question text information using a model consistent with the language model to obtain a vector form, concatenate the diffusion offset vector with the vector form to obtain a three-dimensional semantic vector containing the user's question text information and recall information encoding information, and input the large language model for inference to obtain the final output result.
[0011] In some specific embodiments, S1 specifically includes: obtaining the user's question text information Q = {q1, q2, ..., q n}, and encode it using the BGE vector model to obtain the semantic vector E = BGE (Q) = {e1, e2, ..., e n}, the cosine similarity is calculated as Among them, d n It is the encoding form of the text information in the vector database, and the encoding form is obtained by encoding the text information in the vector database.
[0012] In some specific embodiments, S2 specifically includes: obtaining the entity sequence after definition where n n It is the sequence of entity information contained in a single piece of information; the extracted entity information is indexed from the graph database The storage format of the graph database is a quaternion, that is, m = (n n , r n , p n, e n}, where n, r, and p represent entity information 1, the relationship between entities, and entity information 2 respectively; e represents the vector representation of the text information between node n and node p; according to the relationship between entities, the encoding vector corresponding to the additional semantic information and the sentence information is obtained; the encoding vector is used as known information to obtain the entity enhanced graph database information Merge the graph database information G with the information in the vector database to obtain the final background information
[0013] In some specific embodiments, a diffusion model is used in S3 to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model. The diffusion model diffuses the final background information in multiple steps through a parameter matrix, and controls the diffusion stability through KL divergence and multi-head attention mechanism: in, is the diffusion offset output at the previous moment, For approximation distribution.
[0014] In some specific embodiments, during the diffusion process in S3, hyperparameters α, β, and γ are introduced to limit the amplitude of each diffusion step, where α+β+γ=1; a scaling operation is performed on the vector output by the residual structure in the diffusion model, and the expression of the scaling operation is as follows: Among them, χ is the vector output by the FFW module in the diffusion model, X t Is the vector output after the Scale operation; X t Input into the FFN network structure, and perform α offset on the vector output by the FFN network structure to obtain the diffusion offset at this moment.
[0015] In some specific embodiments, the concatenation of the diffusion offset vector and the vector form in S5 is specifically expressed as: Among them, B represents the Batch of the input model, Seq represents the length received by the model, and Dim represents the dimension of the word vector. represents the final diffusion offset vector, E * A vector representing the encoded user information.
[0016] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.
[0017] According to a third aspect of the present invention, a large model retrieval enhancement generation system based on diffusion alignment is proposed, comprising:
[0018] An information extraction unit is configured to obtain the user's question text information, encode it using a vector model to obtain a semantic vector, calculate the cosine similarity between the semantic vector and the vectors obtained in sequence from the vector database, and extract text information and vector information whose similarity meets a threshold;
[0019] A background information acquisition unit is configured to perform entity extraction on the text information, define entity information categories and make them consistent with the definitions in the graph database, obtain entity sequences, index the extracted entity information in the graph database, obtain additional semantic information and corresponding encoding vectors according to entities and entity relationships, use the encoding vectors as known information to obtain entity-enhanced graph database information, and then merge the graph database information with the information in the vector database to obtain final background information;
[0020] A diffusion offset unit is configured to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model, introduce preset parameters and calculate to obtain a diffusion offset vector; continue to input the diffusion offset vector and the updated vector into the semantic diffusion structure, repeat the diffusion operation to obtain the information diffusion offset vector, subtract the diffusion offset vector from the information diffusion offset vector bit by bit to obtain a semantic vector that eliminates pseudo-semantic noise, and repeat until the number of steps reaches a preset value to obtain the final diffusion offset vector;
[0021] The result output unit is configured to encode the user's question text information using a model consistent with the language model to obtain a vector form, concatenate the diffusion offset vector with the vector form, obtain a three-dimensional semantic vector containing the user's question text information and recall information encoding information, and input the large language model for inference to obtain the final output result.
[0022] In some specific embodiments, the information extraction unit specifically includes obtaining the user's question text information Q = {q1, q2, .., q n}, and encode it using the BGE vector model to obtain the semantic vector E = BGE (Q) = {e1, e2, ..., e n}, the cosine similarity is calculated as Among them, d n It is the encoding form of the text information in the vector database, and the encoding form is obtained by encoding the text information in the vector database.
[0023] In some specific embodiments, the entity sequence defined in the background information acquisition unit is where n n It is the sequence of entity information contained in a single piece of information; the extracted entity information is indexed from the graph database The storage format of the graph database is a quaternion, that is, m = (n n , r n , pn , e n ), where n, r, and p represent entity information 1, the relationship between entities, and entity information 2 respectively; e represents the vector representation of the text information between node n and node p; according to the relationship between entities, the encoding vector corresponding to the additional semantic information and the sentence information is obtained; the encoding vector is used as known information to obtain the entity enhanced graph database information Merge the graph database information G with the information in the vector database to obtain the final background information
[0024] In some specific embodiments, a diffusion model is used in the diffusion offset unit to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model. The diffusion model diffuses the final background information in multiple steps through a parameter matrix, and controls the diffusion stability through KL divergence and multi-head attention mechanism: in, is the diffusion offset output at the previous moment, For approximation distribution; in the diffusion process of the diffusion offset unit, hyperparameters α, β, and γ are introduced to limit the amplitude of each step of diffusion, where α+β+γ=1; the vector output by the residual structure in the diffusion model is scaled, and the expression of the scaling operation is as follows Among them, χ is the vector output by the FFW module in the diffusion model, X t Is the vector output after the Scale operation; X t Input into the FFN network structure, and perform α offset on the vector output by the FFN network structure to obtain the diffusion offset at this moment.
[0025] In some specific embodiments, the result output unit concatenates the diffusion offset vector with the vector form as follows: Among them, B represents the Batch of the input model, Seq represents the length received by the model, and Dim represents the dimension of the word vector. represents the final diffusion offset vector, E * A vector representing the encoded user information.
[0026] The present invention proposes a large-model retrieval enhancement generation method and system based on diffusion alignment, which maps the semantic distribution of the vector retrieval model into the language model, and then embeds the semantic features of the vector into the encoded text vector, avoiding the re-encoding of the text and greatly improving the reasoning speed of the model. At the same time, the invention combines entity recognition technology and graph database indexing, greatly improving the breadth and accuracy of recall information and reducing the illusion, toxicity and instability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and together with the description are used to explain the principles of the present invention. Other embodiments and many expected advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. Other features, objects and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments made with reference to the following drawings:
[0028] Figure 1 is a flow chart of a large model retrieval enhancement generation method based on diffusion alignment according to an embodiment of the present application;
[0029] Figure 2 It is a flowchart of a large model retrieval enhancement generation method based on diffusion alignment in a specific embodiment of the present application;
[0030] Figure 3 It is a retrieval structure diagram of a specific embodiment of the present application;
[0031] Figure 4 It is a distribution diffusion framework diagram of a specific embodiment of the present application;
[0032] Figure 5 is a diffusion model structure diagram of a specific embodiment of the present application;
[0033] Figure 6 is a schematic diagram of diffusion semantic embedding of a specific embodiment of the present application;
[0034] Figure 7 is a graph of semantic distribution changes during the diffusion model training process of a specific embodiment of the present application;
[0035] Figure 8 This is an architecture diagram of a large model retrieval enhancement generation system based on diffusion alignment according to an embodiment of the present application;
[0036] Fig. 9 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION
[0037] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] Figure 1 FIG. 4 shows a flow chart of a large model retrieval enhancement generation method based on diffusion alignment according to an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:
[0040] S1: Obtain the user's question text information, encode it using a vector model to obtain a semantic vector, calculate the cosine similarity between the semantic vector and the vectors obtained from the vector database in sequence, and extract the text information and vector information whose similarity meets the threshold.
[0041] In a specific embodiment, the user's question text information Q {q1, q2, .., q n}, and encode it using the BGE vector model to obtain the semantic vector E = BGE (Q) = {e1, e2, ..., e n}, the cosine similarity is calculated as Among them, d n It is the encoding form of the text information in the vector database, and the encoding form is obtained by encoding the text information in the vector database.
[0042] S2: Perform entity extraction on text information, define entity information categories and make them consistent with the definitions in the graph database, obtain entity sequences, index the extracted entity information in the graph database, obtain additional semantic information and corresponding encoding vectors based on entities and entity relationships, use the encoding vectors as known information to obtain entity-enhanced graph database information, and then merge the graph database information with the information in the vector database to obtain the final background information.
[0043] In a specific embodiment, the entity sequence after definition is obtained where n n It is the sequence of entity information contained in a single piece of information; the extracted entity information is indexed from the graph database The storage format of the graph database is a quaternion, that is, m = (n n , r n , p n , e n}, where n, r, and p represent entity information 1, the relationship between entities, and entity information 2 respectively; e represents the vector representation of the text information between node n and node p; according to the relationship between entities, the encoding vector corresponding to the additional semantic information and the sentence information is obtained; the encoding vector is used as known information to obtain the entity enhanced graph database information Merge the graph database information G with the information in the vector database to obtain the final background information
[0044] S3: Diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model, introduce preset parameters and calculate the diffusion offset vector.
[0045] In a specific embodiment, a diffusion model is used to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model. The diffusion model diffuses the final background information in multiple steps through a parameter matrix, and controls the diffusion stability through KL divergence and multi-head attention mechanism: in, is the diffusion offset output at the previous moment, For approximation distribution.
[0046] S4: Continue to input the diffusion offset vector and the updated vector into the semantic diffusion structure, repeat S3 to obtain the information diffusion offset vector, subtract the diffusion offset vector from the information diffusion offset vector bit by bit to obtain the semantic vector that eliminates pseudo-semantic noise, and loop until the number of steps reaches the preset value to obtain the final diffusion offset vector.
[0047] In a specific embodiment, in the above diffusion process, hyperparameters α, β, and γ are introduced to limit the amplitude of each diffusion step, where α+β+γ=1; a scaling operation is performed on the vector output by the residual structure in the diffusion model, and the expression of the scaling operation is as follows: Among them, χ is the vector output by the FFW module in the diffusion model, X t Is the vector output after the Scale operation; X t Input into the FFN network structure, and perform α offset on the vector output by the FFN network structure to obtain the diffusion offset at this moment.
[0048] S5: Encode the user's question text information using a model consistent with the language model to obtain a vector form, concatenate the diffusion offset vector with the vector form to obtain a three-dimensional semantic vector containing the user's question text information and recall information encoding information, and input the large language model for inference to obtain the final output result.
[0049] In a specific embodiment, the diffusion offset vector is concatenated with the vector form as follows: Among them, B represents the Batch of the input model, Seq represents the length received by the model, and Sim represents the dimension of the word vector. represents the final diffusion offset vector, E * A vector representing the encoded user information.
[0050] In a specific embodiment, Figure 2 A flowchart of a large model retrieval enhancement generation method based on diffusion alignment according to a specific embodiment of the application is shown. Figure 2 As shown, the following steps are included:
[0051] Step 1: A natural language question raised by the user; use retrieval and recall technology to obtain a vector representation of the relevant text information.
[0052] In a specific embodiment, Figure 3 A retrieval structure diagram of a specific embodiment of the present application is shown as follows: Figure 3 As shown in the figure, the specific process of retrieval and recall in the vector database includes:
[0053] After the user input text is encoded into a semantic vector, the similarity is calculated with the stored vectors in the database.
[0054] The vectors that meet the conditions are filtered out through the set similarity threshold.
[0055] The filtered vectors represent the background information most relevant to the user's input question and are used as input for the next step of processing.
[0056] In a specific embodiment, the user's question (text information) Q={q1, q2, .., q n The BGE vector model is used to encode the text information, and the encoded semantic vector E = BGE (Q) = {e1, e2, ..., e n}. After obtaining the encoded semantic vector, vectors are obtained from the vector database in sequence, and the similarity is calculated with the encoded semantic vector. The calculation method is cosine similarity:
[0057] where d n It is the encoding form of the text information in the vector database, which is obtained by encoding the text information in the vector database T = {t1, t2, ..., t n}, D=BGE(T)={d1, d2,...,d n After calculating the similarity between the user text semantics and the background knowledge in the database, extract the text information and vector information whose similarity meets the threshold. Among them, σ is the preset threshold, when D recall When the threshold requirement is met, the second step of entity extraction is performed, otherwise the next piece of information is obtained.
[0058] Step 2: Extract entities from text information to obtain entity sequences and index them into the graph database to obtain quadruple sequences. Obtain information from the graph database and merge it with the retrieved recall information.
[0059] In a specific embodiment, the text of the relevant background information obtained in the first step is subjected to entity extraction to define the entity information to be identified. The entity information category defined needs to be consistent with the definition in the graph database. where n n It is the sequence of entity information contained in a single piece of information. The extracted entity information is indexed from the graph database The storage format of the graph database is a four-tuple, that is, m = (n n , r n , p n , e n ), where n, r, and p represent entity information 1, the relationship between entities, and entity information 2 respectively; e represents the vector representation of the text information between node n and node p. According to the relationship between entities, the encoding vector corresponding to the additional semantic information and the sentence information is obtained; the encoding vector is used as known information to obtain the entity-enhanced graph database information G: Merge the graph database information G with the information in the vector database to obtain the final background information
[0060] In a specific embodiment, Figure 7 FIG. 4 shows a diagram of semantic distribution changes during the diffusion model training process of a specific embodiment of the present application, such as Figure 7 As shown in the figure, a to c represent the process in which the model semantic distribution gradually diffuses into the vector semantic distribution during the training process, where the difference between the vector semantic distribution encoded by the retrieval model and the semantic distribution of the language model can be regarded as semantic noise. During the training process, the vector semantic distribution gradually diffuses into the model semantic distribution, thereby avoiding the problem of secondary encoding of the model by text information.
[0061] Step 3: The semantic diffusion model diffuses relevant information from the vector semantic space to the large model semantic space.
[0062] In a specific embodiment, the merged vector information is diffused from the original BGE word vector space to the word vector space of the language model: Among them, t is a hyperparameter set artificially, indicating the number of steps in the diffusion process. is the background information in the form of vector representation obtained in step 3. The specific diffusion model is as follows Figure 5 As shown: where coefficient matrix is the parameter matrix, t=1 is set as the initial value, t is one-hot encoded, and the encoded vector representation is obtained k is the total number of diffusion steps. After multiplication, we get the pseudo semantic noise vector corresponding to time t. In order to facilitate subsequent embedding into the vector representation of the semantic model, we need to map the k-dimensional vector to the same dimension dim as the model: Figure 5 The Token Embedding in step 2 is the vector representation of the background information. In order to ensure the stability of the diffusion process, the present invention introduces three preset hyperparameters (α, β, γ) to limit the amplitude of each diffusion step, where α+β+γ=1. The FFW module is designed to further constrain the diffusion offset, and its specific operation is as follows: To ensure the effectiveness of the diffusion process, it is necessary to ensure that the newly obtained diffusion offset vector is similar to the diffusion offset vector at the previous time point during the diffusion process, and use KL divergence to measure the difference between the new diffusion offset vector obtained each time and the diffusion offset vector obtained in the previous step: in, is the diffusion offset vector output at the previous moment, To approximate When the approximation is less than the set threshold, it is necessary to re-calibrate Do regularization, It is the vector output after the multi-head attention mechanism.
[0063] In a specific embodiment, Scale represents a scaling operation on the vector output by the residual structure, and the specific expression is as follows: Among them, χ is the vector after FFW output, X t It is the vector output after the Scale operation. After the Scale operation, X t The variance of the distribution does not change, and the mean becomes the original Get X t Then input it into the FFN network structure and get Factor Shift means to perform α shift on the vector information output by FFN. The specific expression is as follows: After Factor Shift, the diffusion offset vector at this moment is obtained, and its distribution can be approximated by normal distribution: Among them, μ t ∝α,V t-1 ∝β.
[0064] In a specific embodiment, the diffusion offset vector V obtained above is n} and the updated T=t+1 continue to be input into the semantic diffusion structure (such as Figure 4 As shown), repeat the above operation to obtain the information diffusion offset vector The semantic vector that eliminates pseudo-semantic noise is obtained by subtracting the diffusion offset vector V obtained in the previous step from the offset vector obtained in this diffusion step bit by bit. Repeat this operation until the number of steps reaches the set value, that is, T = k. When the cycle reaches the specified number of steps, the final diffusion offset vector is obtained.
[0065] Step 4: The large language model encodes the user's question and outputs the word vector; concatenates the vector representation of the semantic space and the user question encoding; and embeds the large model for forward propagation.
[0066] In a specific embodiment, the user information Q in the first step is text-encoded, and the adopted Embedding model is consistent with the language model (Llama is used as an example in this application, but this method can be applied to any decoder-only large language model). The vector form of the encoded user information is obtained The diffusion offset vector obtained in the third step is combined with E * To splice (such as Figure 6 As shown), a concatenated vector representation is obtained, which contains both the user's text information and the encoding information of the recalled information without encoding the recalled text information. Moreover, the vector representation is semantically mapped by the diffusion model in step three and is consistent with the distribution of the large language model in the semantic space. Figure 6 The Text-Embedding in this step is the E * , K-Embedding is the final diffusion offset vector obtained in the fourth step The concatenated expression is Among them, B represents the batch of input model, Seq represents the length received by the model, and Dim represents the dimension of word vector. In particular, the concatenation method is to concatenate from the Seq dimension, and the concatenated vector is a three-dimensional semantic vector. Get the concatenated vector Finally, the final output result is obtained through reasoning of the large language model.
[0067] The present invention designs a large language model retrieval embedding generation method based on diffusion alignment for large language model retrieval recall generation RAG. Since the inconsistency in semantic distribution between the vector model and the language model is caused by the error in the model itself, this error is closely related to the pre-training data and model structure of the model. Based on this assumption, the application trains a semantic diffusion model to fill this error, maps the semantic distribution of the vector retrieval model into the language model, and then embeds the semantic features of the vector into the encoded text vector, avoiding the re-encoding of the text and greatly improving the reasoning speed of the model. At the same time, the invention combines entity recognition technology and graph database indexing, which greatly improves the breadth and accuracy of recall information and reduces the illusion, toxicity and instability of the model.
[0068] Figure 8 FIG. 1 shows an architecture diagram of a large model retrieval enhancement generation system based on diffusion alignment according to an embodiment of the present application. Figure 8 The system includes an information extraction unit 801, a background information acquisition unit 802, a diffusion offset unit 803 and a result output unit 804, wherein the information extraction unit 801 is configured to obtain the user's question text information, encode it using a vector model to obtain a semantic vector, perform cosine similarity calculation on the semantic vector and the vectors obtained in sequence from the vector database, and extract text information and vector information whose similarity meets the threshold; the background information acquisition unit 802 is configured to perform entity extraction on the text information, define the entity information category and make it consistent with the definition in the graph database, obtain an entity sequence, index the extracted entity information in the graph database, obtain additional semantic information and corresponding encoding vectors according to entities and entity relationships, use the encoding vector as known information to obtain entity-enhanced graph database information, and then compare the graph database information with the information in the vector database The final background information is obtained by merging; the diffusion offset unit 803 is configured to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model, introduce preset parameters and calculate the diffusion offset vector; the diffusion offset vector and the updated vector are continuously input into the semantic diffusion structure, the diffusion operation is repeated to obtain the information diffusion offset vector, the diffusion offset vector is subtracted bit by bit from the information diffusion offset vector to obtain the semantic vector that eliminates pseudo-semantic noise, and the cycle is repeated until the number of steps reaches the preset value to obtain the final diffusion offset vector; the result output unit 804 is configured to use a model consistent with the language model to perform text encoding on the user's question text information to obtain a vector form, the diffusion offset vector is concatenated with the vector form to obtain a three-dimensional semantic vector containing the user's question text information and the recall information encoding information, and the large language model is input for inference to obtain the final output result.
[0069] Reference below Fig. 9, which shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0070] like Fig. 9 As shown, the computer system includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0071] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed.
[0072] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0073] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0074] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0075] The modules involved in the embodiments of the present application may be implemented by software or by hardware.
[0076] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device: obtains the text information input by the user, and encodes the text information using a vector model to obtain the encoded semantic vector; retrieves text information similar to the encoded semantic vector from the vector database, and selects qualified text information according to the similarity threshold; performs entity extraction on the qualified text information, indexes the extracted entity information with the entity relationship stored in the graph database, and generates enhanced background information containing additional semantic information; diffuses the enhanced background information from the original word vector space to the word vector space of the large language model, and obtains a diffuse offset vector consistent with the semantic distribution of the language model; splices the diffuse offset vector with the encoded vector of the user input text information, inputs the large language model for reasoning, and outputs a generated result containing retrieval enhancement information.
[0077] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A large model retrieval enhancement generation method based on diffusion alignment, characterized in that: include: S1: Obtain the user's question text information, encode it using a vector model to obtain a semantic vector, calculate the cosine similarity between the semantic vector and the vectors obtained in sequence from the vector database, and extract the text information and vector information whose similarity meets the threshold; S2: extracting entities from the text information, defining entity information categories and making them consistent with the definitions in the graph database, obtaining entity sequences, indexing the extracted entity information in the graph database, obtaining additional semantic information and corresponding encoding vectors based on entities and entity relationships, using the encoding vectors as known information to obtain entity-enhanced graph database information, and then merging the graph database information with the information in the vector database to obtain final background information; S3: diffusing the final background vector information from the original BGE word vector space to the word vector space of the language model, introducing preset parameters and calculating a diffusion offset vector; S4: continue to input the diffusion offset vector and the updated vector into the semantic diffusion structure, repeat S3 to obtain the information diffusion offset vector, subtract the diffusion offset vector from the information diffusion offset vector bit by bit to obtain the semantic vector that eliminates the pseudo-semantic noise, and repeat until the number of steps reaches a preset value to obtain the final diffusion offset vector; S5: The user's question text information is encoded using a model consistent with the language model to obtain a vector form, the diffusion offset vector is concatenated with the vector form to obtain a three-dimensional semantic vector containing the user's question text information and recall information encoding information, and the vector is input into the large language model for inference to obtain the final output result.
2. The large model retrieval enhancement generation method based on diffusion alignment according to claim 1 is characterized in that: The S1 specifically includes: obtaining the user's question text information Q={q1, q2, .., q n }, and encode it using the BGE vector model to obtain the semantic vector E = BGE (Q) = {e1, e2, ..., e n }, the cosine similarity is calculated as Among them, d n It is the encoding form of the text information in the vector database, and the encoding form is obtained by encoding the vector library text information.
3. The large model retrieval enhancement generation method based on diffusion alignment according to claim 1 is characterized in that: The S2 specifically includes: obtaining the entity sequence after definition Where n n is a sequence of entity information contained in a single piece of information; the extracted entity information is indexed from the graph database The storage format of the graph database is four-tuple, namely Where n, r, and p represent entity information 1, the relationship between entities, and entity information 2, respectively; Represents the vector representation of the text information between node n and node p; Based on the relationship between entities, obtain the encoding vector corresponding to the additional semantic information and the sentence information; Take the encoding vector as known information to obtain the entity enhanced graph database information Merge the graph database information G with the information in the vector database to obtain the final background information 4. The large model retrieval enhancement generation method based on diffusion alignment according to claim 3 is characterized in that: In S3, the diffusion model is used to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model. The diffusion model diffuses the final background information in multiple steps through a parameter matrix, and controls the diffusion stability through KL divergence and multi-head attention mechanism: in, is the diffusion offset output at the previous moment, For approximation distribution.
5. The large model retrieval enhancement generation method based on diffusion alignment according to claim 4 is characterized in that: In the diffusion process of S3, hyperparameters α, β, and γ are introduced to limit the amplitude of each diffusion step, where α+β+γ=1; the vector output by the residual structure in the diffusion model is scaled, and the expression of the scaling operation is as follows: Where χ is the vector output by the FFW module in the diffusion model, X t Is the vector output after the Scale operation; X t The vector output by the FFN network structure is inputted, and an α offset is performed on the vector outputted by the FFN network structure to obtain the diffusion offset at this moment.
6. The large model retrieval enhancement generation method based on diffusion alignment according to claim 5 is characterized in that: In S5, the concatenation of the diffusion offset vector and the vector form is specifically expressed as: Among them, B represents the Batch of the input model, Seq represents the length received by the model, and Dim represents the dimension of the word vector. represents the final diffusion offset vector, E * A vector representing the encoded user information.
7. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 6 is implemented.
8. A large model retrieval enhancement generation system based on diffusion alignment, characterized in that: include: An information extraction unit is configured to obtain the user's question text information, encode it using a vector model to obtain a semantic vector, calculate the cosine similarity between the semantic vector and the vectors obtained in sequence from the vector database, and extract text information and vector information whose similarity meets a threshold; A background information acquisition unit is configured to perform entity extraction on the text information, define entity information categories and make them consistent with the definitions in the graph database, obtain entity sequences, index the extracted entity information in the graph database, obtain additional semantic information and corresponding encoding vectors according to entities and entity relationships, use the encoding vectors as known information to obtain entity-enhanced graph database information, and then merge the graph database information with information in the vector database to obtain final background information; A diffusion offset unit is configured to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model, introduce preset parameters and calculate a diffusion offset vector; continue to input the diffusion offset vector and the updated vector into the semantic diffusion structure, repeat the diffusion operation to obtain an information diffusion offset vector, subtract the diffusion offset vector from the information diffusion offset vector bit by bit to obtain a semantic vector that eliminates pseudo-semantic noise, and repeat until the number of steps reaches a preset value to obtain a final diffusion offset vector; The result output unit is configured to perform text encoding on the user's question text information using a model consistent with the language model to obtain a vector form, concatenate the diffusion offset vector with the vector form to obtain a three-dimensional semantic vector containing the user's question text information and recall information encoding information, and input the vector into the large language model for inference to obtain the final output result.
9. The large model retrieval enhancement generation system based on diffusion alignment according to claim 8 is characterized in that: The information extraction unit specifically includes obtaining the user's question text information Q={q1, q2, .., q n }, and encode it using the BGE vector model to obtain the semantic vector E = BGE (Q) = {e1, e2, ..., e n }, the cosine similarity is calculated as Among them, d n It is the encoding form of the text information in the vector database, and the encoding form is obtained by encoding the vector library text information.
10. The large model retrieval enhancement generation system based on diffusion alignment according to claim 8 is characterized in that: The entity sequence defined in the background information acquisition unit Where n n is a sequence of entity information contained in a single piece of information; the extracted entity information is indexed from the graph database The storage format of the graph database is four-tuple, namely Where n, r, and p represent entity information 1, the relationship between entities, and entity information 2, respectively; Represents the vector representation of the text information between node n and node p; Based on the relationship between entities, obtain the encoding vector corresponding to the additional semantic information and the sentence information; Take the encoding vector as known information to obtain the entity enhanced graph database information Merge the graph database information G with the information in the vector database to obtain the final background information 11. The large model retrieval enhancement generation system based on diffusion alignment according to claim 10 is characterized in that: The diffusion offset unit uses a diffusion model to diffuse the final background vector information from the original BGE word vector space to the word vector space of the language model. The diffusion model diffuses the final background information in multiple steps through a parameter matrix, and controls the diffusion stability through KL divergence and multi-head attention mechanism: in, is the diffusion offset output at the previous moment, For approximation distribution; in the diffusion process of the diffusion offset unit, hyperparameters α, β, and γ are introduced to limit the amplitude of each diffusion step, where α+β+γ=1; a scaling operation is performed on the vector output by the residual structure in the diffusion model, and the expression of the scaling operation is as follows Where χ is the vector output by the FFW module in the diffusion model, X t Is the vector output after the Scale operation; X t The vector output by the FFN network structure is inputted, and an α offset is performed on the vector outputted by the FFN network structure to obtain the diffusion offset at this moment.
12. The large model retrieval enhancement generation system based on diffusion alignment according to claim 11 is characterized in that: The result output unit splices the diffusion offset vector with the vector form as specifically represented as follows: Among them, B represents the Batch of the input model, Seq represents the length received by the model, and Dim represents the dimension of the word vector. represents the final diffusion offset vector, E * A vector representing the encoded user information.
Citation Information
Patent Citations
Patent document query method based on diffusion model and computer equipment
CN115794999A
Online intelligent question answering method and device based on instruction fine tuning and retrieval enhancement generation
CN117688163A
Retrieval method and system based on domain-enhanced large language model
CN118796978A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Cited By
Three-dimensional content generation method and device based on retrieval enhancement and personalized generation
CN121392164A
Large model retrieval enhancement generation system based on vector database
CN122346469A