Intelligent question answering method based on diffusion model and large language model

By introducing a text diffusion model into the intelligent Q&A system to extract and integrate keyword information, the problem of insufficient attention when dealing with complex problems is solved, and the user Q&A experience is significantly improved.

CN120045680APending Publication Date: 2025-05-27北京粉笔上岸科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328636.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing intelligent Q&A method based on large language model is difficult to pay attention to keywords correctly when dealing with complex problems, resulting in failure to understand user questions and answering non-questions, and bringing users a bad Q&A experience.

Method used

Using intelligent Q&A based on diffusion model and large language model, the text diffusion model is trained to extract keyword information and integrate it into the text embedding layer of the large language model to enhance the model's attention to key parts of complex problems.

Benefits of technology

It effectively enhances the attention of the large language model to key parts of complex problems, helps the model better understands the user's complex problems and improves the user's Q&A experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045680A_ABST
    Figure CN120045680A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent question answering method based on a diffusion model and a large language model, and relates to the technical field of question answering models, and the method comprises the following steps: collecting a large number of user questions, texts and keywords corresponding to the questions and texts, and training a text diffusion model based on the questions and texts; collecting a large amount of book texts and user-teacher question and answer data, pre-training an open-source large language model by using the book texts, and adjusting the large language model by using the user-teacher question and answer data; after a user inquires a question, the text diffusion model converts the user question into a keyword vector, the keyword vector and the user question are input into the large language model together, the attention degree of each character in the question is calculated, and question answering content is output; according to the method, a text diffusion model is trained to extract keyword information, and the keyword information is fused into a text embedding layer of a large language model, so that the keyword information is fused into a text embedding vector of the large language model, and the attention degree of the model on key parts in complex problems is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of question - answering models, and particularly to an intelligent question - answering method based on diffusion models and large - language models. Background Art

[0002] The intelligent question - answering method based on large - language models refers to using a large - language model trained on a large amount of text. According to the question input by the user, the "attention module" in the model structure is used to automatically calculate the degree of attention to each word in the question, so as to generate a corresponding answer.

[0003] In the prior art, although large - language models can quickly answer the questions input by users, limited by the computing power of the "attention module", when the user's question is relatively complex, its attention mechanism is difficult to correctly focus on the corresponding keywords, and large - language models are prone to not understanding the user's question and answering irrelevantly, thus bringing a bad question - answering experience to users. Therefore, the present invention proposes an intelligent question - answering method based on diffusion models and large - language models to solve the problems existing in the prior art. Summary of the Invention

[0004] In view of the above problems, the present invention proposes an intelligent question - answering method based on diffusion models and large - language models. This intelligent question - answering method based on diffusion models and large - language models trains a text diffusion model to extract keyword information and integrates it into the text embedding layer of the large - language model, thereby integrating keyword information into the text embedding vector of the large - language model and enhancing the model's attention to the key parts in complex questions.

[0005] To achieve the object of the present invention, the present invention is realized through the following technical solutions: An intelligent question - answering method based on diffusion models and large - language models, comprising the following steps:

[0006] S1: Collect a large number of user questions, texts, and their corresponding keywords, and train a text diffusion model based on this.

[0007] S2: Collect a large number of book texts and user - teacher Q&A data, pre - train an open - source large - language model using the book texts, and adjust the large - language model using the user - teacher Q&A data.

[0008] S3: After the user asks a question, the text diffusion model converts the user's question into a keyword vector, inputs it together with the user's question into the large - language model, calculates the degree of attention to each word in the question, and outputs the question - answering content.

[0009] Further improvement lies in: In the above S1, a large number of user questions, texts, and their corresponding keywords are training data K, the text is text, the keyword is keywords, and training the text diffusion model includes the following processes:

[0010] For a piece of training data (text, keyword_1, keyword_2, …, keyword_n) in K, where n is the number of keywords corresponding to text, first use the BERT model to embed this training data into the vector space, obtaining a vector group E = [E_t, E_k1, E_k2, …, E_kn], where E_t is the embedding vector of text and E_ki is the embedding vector of the i-th keyword;

[0011] Use Markov transformation to convert E = [E_t, E_k1, E_k2, …, E_kn] into an initial random variable X_0 = [x0_t, x0_k1, x0_k2, …, x0_kn] that follows a normal distribution. The transformation formula is as follows:

[0012] x0_t ∼ N(E_t, a * I), x0_k1 ∼ N(E_k1, a * I), …

[0013] where I is a vector of all 1s, a is a scaling factor between 0 and 1, N(u, s) is a normal distribution with mean u and variance s, and ∼ is the sampling operation;

[0014] Add random noise that follows the N(0, a) distribution to the vectors x0_k1, x0_k2, …, x0_kn in X_0 successively for t times, where t is randomly selected from 1 to 100. When adding random noise for the first time, the random variable X_0 is converted to X_1, and so on:

[0015] X_1 = [x0_t, x1_k1, x1_k2, …, x1_kn]

[0016] where x1_k1 ∼ N(\sqrt{1 - a} * x0_k1, a * I), x2_k1 ∼ N(\sqrt{1 - a} * x0_k2, a * I), …;

[0017] Use the Transformer Encoder model to be trained to restore the variable X_t = [x0_t, xt_k1, xt_k2, …, xt_kn] with t times of random noise added to E = [E_t, E_k1, E_k2, …, E_kn]. Let the restored E by the Transformer Encoder be E_p. Then the loss function of the diffusion model is:

[0018] L = MSE(E_p, E) + R(X_1)

[0019] where MSE is the mean squared error and R(X_1) is the L2 regularization term.

[0020] A further improvement lies in: using the BERT word embedding model to convert the text and keywords into vectors E t and E k respectively; randomly adding noise from the normal distribution to the keyword vector t times to obtain X k ; inputting E t and X k into the transformer encoder model, and restoring the keyword vector X k after adding noise; calculating the loss value between E k and :

[0021]

[0022] Calculating the gradient of the loss, and using gradient descent to update the transformer encoder model, and continuously repeating the above process.

[0023] A further improvement lies in: in S2, the pre-trained open-source large language model refers to: an open-source large language model that has been pre-trained on a large amount of text and books, and on this basis, continues to use private and new text and books to continue training the model. The training objective is that given the above text, the model can accurately predict the next word; adjusting the large language model refers to: based on the pre-trained open-source large language model, using user-teacher Q&A data to train the model. The training objective is that given the input question, the model can output the corresponding answer.

[0024] A further improvement lies in: in S3, after the user asks a question, the text diffusion model first embeds the user question into a text vector E t using the BERT word embedding model; then randomly samples multiple noises from the normal distribution to obtain X k ; inputting E t and X k into the transformer encoder to obtain the restored keyword vector E k , which specifically includes the following steps:

[0025] When the user inputs a question Q, use BERT to encode Q into a vector E_Q, and use Markov transformation to convert E_Q into a random variable x0_Q;

[0026] Randomly sample n random vectors [xt_k1, xt_k2,..., xt_kn] from the N(0, I) distribution, and concatenate them with x0_Q to form X_t = [x0_Q, xt_k1, xt_k2,..., xt_kn].

[0027] Use the Transformer Encoder model to restore X_t to E = [E_Q, E_k1, E_k2,..., E_kn], where [E_k1, E_k2,..., E_kn] are the vectors of the n keywords corresponding to Q;

[0028] At this time, the keyword information of the user input question Q is encoded into the vector group [E_k1, E_k2,..., E_kn] by the text diffusion model.

[0029] A further improvement is that in S3, E k and the user question are input into the large language model together, where E k will be inserted into the corresponding token vectors in the attention calculation of the large language model, specifically including the following steps:

[0030] Before the large language model generates the answer to Q, it uses the embedding layer in its structure to encode the question Q into a token vector group [e_q1, e_q2, e_q3,..., e_qL], where e_qi is the vector embedding of the i-th token in Q;

[0031] Add the vector group [E_k1, E_k2,..., E_kn] with keyword information to the corresponding token vectors in the token vector group. Suppose n = 1, and E_k1 corresponds to the vector representation of the keyword composed of the 2nd and 3rd tokens, then each of the token vectors e_q2 and e_q3 will be added with E_k1:

[0032] new_e_q2 = e_q2 + E_k1; new_e_q3 = e_q3 + E_k

[0033] Thus, the token vector group of the large language model is [e_q1, new_e_q2, new_e_q3,..., e_qL].

[0034] A further improvement is that in S3, calculate the attention degree for each word in the question and output the answer content, including the following steps:

[0035] The large language model transmits the token vector group after adding keyword information to the corresponding "attention module" in its structure, so as to calculate the attention degree for each word in the question;

[0036] The vectors corresponding to the keywords in the token vector group incorporate keyword information, thereby enhancing the attention degree of the model to the key parts in the question Q and outputting the answer content.

[0037] The beneficial effects of the present invention are:

[0038] 1. The present invention trains a text diffusion model to extract keyword information and integrates it into the text embedding layer of a large language model, thereby integrating keyword information into the text embedding vectors of the large language model and enhancing the model's attention to key parts in complex questions.

[0039] 2. The present invention uses keyword vector information to enhance the large language model's attention to key parts in complex questions, thereby helping the model better understand the user's complex questions and improving the user's question-answering experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of the training text diffusion model of the present invention;

[0041] Figure 2 is a flowchart of the training large language model of the present invention;

[0042] Figure 3 is a flowchart of the online question and answer of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0043] To deepen the understanding of the present invention, the following will further elaborate on the present invention in conjunction with embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.

[0044] Embodiment 1

[0045] According to Figure 1 、 2 、and Figure 3, this embodiment proposes an intelligent question-answering method based on a diffusion model and a large language model, including the following steps: training a text diffusion model, training a large language model, and online question and answer.

[0046] Training the text diffusion model:

[0047] First, collect a large number of user questions, texts, and their corresponding keywords. Then, train the text diffusion model through the following process:

[0048] Use the BERT word embedding model to convert the text and keywords into vectors E t 、E k .

[0049] Randomly add t times of normal distribution noise to the keyword vector to obtain X k

[0050] Input E t 、X k into the transformer encoder model, and restore the keyword vector X k after adding noise to Calculate E k and The loss value is:

[0051]

[0052] Calculate the gradient of the loss, use gradient descent to update the transformer encoder model, and repeat the above process.

[0053] The principle of extracting keyword vectors in the text diffusion model is: during training, noise is continuously added to the keyword vectors to make them tend more and more towards a certain distribution (such as normal distribution), and then a restorer (transformer encoder model) is trained. The restorer removes noise and restores the keyword vectors after adding noise based on the text vectors and the keyword vectors after adding noise. Later, during reasoning, when new text is available, the distributed noise is directly restored to the keyword vector based on the vector of the new text and the randomly sampled distributed noise.

[0054] Training a large language model:

[0055] Collect a large amount of book text and user-teacher question-and-answer data;

[0056] Use book text to pre-train an open-source large language model;

[0057] Tuning large language models using user-teacher question-answer data.

[0058] Pre-trained open source large language model refers to: based on the open source large language model that has been pre-trained on a large number of texts and books, the model is continued to be trained using private, new texts and books. The training goal is to input the previous text and the model can accurately predict the next word / word; adjusted large language model refers to: based on the pre-trained open source large language model, the model is trained using user-teacher question and answer data. The training goal is to input questions and the model can output corresponding answers.

[0059] Online Q&A:

[0060] After the user asks a question, the text diffusion model first embeds the user question into a text vector E using the BERT word embedding model. t ;

[0061] Then randomly sample multiple noises from a normal distribution to obtain X k ; E t , X k Input into transformer encoder to get the restored keyword vector E k ;

[0062] Finally, E kInput into the large language model together with the user's question, where E k will be inserted into the corresponding token vectors during the attention calculation of the large language model;

[0063] Finally, the model outputs the answer content.

[0064] Example Two

[0065] According to Figure 1 、 2 、as shown in Figure 3, this embodiment proposes an intelligent question-answering method based on a diffusion model and a large language model, including the following steps:

[0066] Collect a large amount of (text, keywords) training data K, where text is the text, which can be any form of written content such as articles, conversations, news, etc., and keywords are one or more keywords that can summarize the corresponding theme of the text.

[0067] Train a text diffusion model on the K data. The training process is as follows:

[0068] For a certain piece of training data (text, keyword_1, keyword_2,..., keyword_n) in K, where n is the number of keywords corresponding to text. First, use the BERT model to embed this training data into the vector space to obtain a vector group E = [E_t, E_k1, E_k2,..., E_kn], where E_t is the embedding vector of text, and E_ki is the embedding vector of the i-th keyword;

[0069] Use Markov transformation to convert E = [E_t, E_k1, E_k2,..., E_kn] into an initial random variable X_0 = [x0_t, x0_k1, x0_k2,..., x0_kn] that follows a normal distribution. The conversion formula is as follows:

[0070] x0_t ∼ N(E_t, a*I), x0_k1 ∼ N(E_k1, a*I),...

[0071] where I is a vector of all 1s, a is a scaling factor between 0 and 1, N(u, s) is a normal distribution with mean u and variance s, and ∼ is the sampling operation;

[0072] Add random noise that follows the N(0, a) distribution to the vectors x0_k1, x0_k2,..., x0_kn in X_0 t times in sequence, where t is randomly selected from 1 to 100. For example, when adding random noise for the first time, the random variable X_0 is converted to X_1:

[0073] X_1 = [x0_t, x1_k1, x1_k2,..., x1_kn]

[0074] where \(x1_{k1}\sim N(\sqrt{1 - a}*x0_{k1}, a*I)\), \(x2_{k1}\sim N(\sqrt{1 - a}*x0_{k2}, a*I)\), \(\cdots\);

[0075] Use the Transformer Encoder model to be trained to restore the variable \(X_t = [x0_t, xt_{k1}, xt_{k2}, \cdots, xt_{kn}]\) with \(t\) times of random noise added to \(E = [E_t, E_{k1}, E_{k2}, \cdots, E_{kn}]\). Assume that \(E\) restored by the Transformer Encoder is \(E_p\), then the loss function of the diffusion model is:

[0076] \(L = MSE(E_p, E)+R(X_1)\)

[0077] where \(MSE\) is the mean squared error and \(R(X_1)\) is the \(L2\) regularization term.

[0078] When the user inputs the question \(Q\), the text diffusion model will generate \(n\) corresponding keyword vectors \([E_{k1}, E_{k2}, \cdots, E_{kn}]\) from \(Q\), and the generation process is as follows:

[0079] First, use BERT to encode \(Q\) into a vector \(E_Q\), and use Markov transformation to convert \(E_Q\) into a random variable \(x0_Q\);

[0080] Randomly sample \(n\) random vectors \([xt_{k1}, xt_{k2}, \cdots, xt_{kn}]\) from the \(N(0, I)\) distribution, and concatenate them with \(x0_Q\) to form \(X_t = [x0_Q, xt_{k1}, xt_{k2}, \cdots, xt_{kn}]\);

[0081] Use the Transformer Encoder model to restore \(X_t\) to \(E = [E_Q, E_{k1}, E_{k2}, \cdots, E_{kn}]\), where \([E_{k1}, E_{k2}, \cdots, E_{kn}]\) are the vectors of the \(n\) keywords corresponding to \(Q\);

[0082] At this time, the keyword information of the user's input question \(Q\) has been encoded into the vector group \([E_{k1}, E_{k2}, \cdots, E_{kn}]\) by the text diffusion model.

[0083] Before the large language model generates the answer to \(Q\), first use the embedding layer in its structure to encode the question \(Q\) into a token (the smallest character unit recognizable by the large language model) vector group \([e_{q1}, e_{q2}, e_{q3}, \cdots, e_{qL}]\), where \(e_{qi}\) is the vector embedding of the \(i\)-th token in \(Q\), and add the vector group \([E_{k1}, E_{k2}, \cdots, E_{kn}]\) with keyword information to the corresponding token vectors in the token vector group;

[0084] Let \(n = 1\). If \(E_{k1}\) corresponds to the vector representation of the keyword composed of the 2nd and 3rd tokens, then for each token vector of \(e_{q2}\) and \(e_{q3}\), \(E_{k1}\) will be added:

[0085] new_e_q2 = e_q2 + E_{k1}; new_e_q3 = e_q3 + E_{k1}

[0086] Thus, the token vector group of the large language model is \([e_{q1}, new_e_q2, new_e_q3, \ldots, e_{qL}]\).

[0087] The large language model transmits the token vector group with the added keyword information to the corresponding "attention module" in its structure, thereby calculating the degree of attention to each word in the question. Since the vectors corresponding to the keywords in the token vector group have incorporated the keyword information, the model's attention to the key parts of question Q is enhanced.

[0088] Embodiment 3

[0089] This embodiment proposes an intelligent question - answering method based on a diffusion model and a large language model, including the following steps: Use a small - scale language model to rewrite the user's question into an easier - to - understand question, and then let the large model answer it.

[0090] This intelligent question - answering method based on a diffusion model and a large language model trains the text diffusion model to extract keyword information and integrates it into the text embedding layer of the large language model, thereby incorporating keyword information into the text embedding vectors of the large language model and enhancing the model's attention to the key parts of complex questions. Moreover, the present invention utilizes keyword vector information to enhance the model's attention to the key parts of complex questions, thereby helping the model better understand the user's complex questions and improving the user's question - answering experience.

[0091] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above - mentioned embodiments. What is described in the above - mentioned embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent question answering method based on a diffusion model and a large language model, characterized in that: The following steps are involved: S1: Collect a large number of user questions, texts and their corresponding keywords, and train the text diffusion model based on them; S2: Collect a large amount of book texts and user-teacher question-and-answer data, use the book texts to pre-train an open-source large language model, and use the user-teacher question-and-answer data to adjust the large language model; S3: After the user asks a question, the text diffusion model converts the user question into a keyword vector, inputs it into the large language model together with the user question, calculates the degree of attention paid to each word in the question, and outputs the answer content.

2. The intelligent question answering method based on diffusion model and large language model according to claim 1, characterized in that: In S1, a large number of user questions, texts and their corresponding keywords are training data K, the text is text, and the keywords are keywords. Training the text diffusion model includes the following process: For a piece of training data in K (text, keyword_1, keyword_2, …, keyword_n), where n is the number of keywords corresponding to text, first use the BERT model to embed the training data into the vector space to obtain a vector group E = [E_t, E_k1, E_k2, …, E_kn], where E_t is the embedding vector of text and E_ki is the embedding vector of the i-th keyword; Use Markov transformation to transform E = [E_t, E_k1, E_k2, ..., E_kn] into an initial random variable X_0 = [x0_t, x0_k1, x0_k2, ..., x0_kn] that obeys a normal distribution. The transformation formula is as follows: x0_t~N(E_t, a*I), x0_k1~N(E_k1, a*I),… Where I is a vector of all 1s, a is a scaling factor between 0 and 1, N(u, s) is a normal distribution with mean u and variance s, and ~ is a sampling operation; Add t times of random noise from N(0, a) distribution to the vectors x0_k1, x0_k2, …, x0_kn in X_0, where t is randomly selected from 1 to 100. When adding random noise for the first time, the random variable X_0 is converted to X_1, and so on: X_1=[x0_t,x1_k1,x1_k2,…,x1_kn] Among them, x1_k1~N(\sqrt{1-a}*x0_k1, a*I), x2_k1~N(\sqrt{1-a}*x0_k2, a*I),…; Use the Transformer Encoder model to be trained to restore the variable X_t = [x0_t, xt_k1, xt_k2, ..., xt_kn] with t times of random noise to E = [E_t, E_k1, E_k2, ..., E_kn]. Suppose E restored by Transformer Encoder is E_p, then the loss function of the diffusion model is: L=MSE(E_p,E)+R(X_1) Among them, MSE is the mean square error and R(X_1) is the L2 regularization term.

3. The intelligent question answering method based on diffusion model and large language model according to claim 2, characterized in that: Use the BERT word embedding model to convert text and keywords into vectors E respectively t 、E k ; Randomly add t times normal distribution noise to the keyword vector to get X k ; E t , X k Input into the transformer encoder model, and add the noise to the keyword vector X k Restore to Calculate E k and The loss value is: Calculate the gradient of the loss, use gradient descent to update the transformer encoder model, and repeat the above process.

4. The intelligent question answering method based on diffusion model and large language model according to claim 1, characterized in that: In S2, pre-training an open source large language model refers to: based on an open source large language model that has been pre-trained on a large number of texts and books, continuing to train the model using private, new texts and books on this basis, and the training goal is to input the previous text, and the model can accurately predict the next word / word; adjusting the large language model refers to: based on the pre-trained open source large language model, using user-teacher question and answer data to train the model, and the training goal is to input questions and the model can output corresponding answers.

5. The intelligent question answering method based on diffusion model and large language model according to claim 3 is characterized by: In S3, after the user asks a question, the text diffusion model first embeds the user question into a text vector E using the BERT word embedding model. t ; Then randomly sample multiple noises from the normal distribution to get X k ; E t , X k Input into transformer encoder to get the restored keyword vector E k , specifically including the following steps: When the user enters a question Q, BERT is used to encode Q into a vector E_Q, and Markov transformation is used to convert E_Q into a random variable x0_Q; Randomly sample n random vectors [xt_k1, xt_k2, …, xt_kn] from N(0, I) distribution and concatenate them with x0_Q to become X_t = [x0_Q, xt_k1, xt_k2, …, xt_kn]. Use the Transformer Encoder model to restore X_t to E = [E_Q, E_k1, E_k2, ..., E_kn], where [E_k1, E_k2, ..., E_kn] is the vector of n keywords corresponding to Q; At this time, the keyword information of the user input question Q is encoded into the vector group [E_k1, E_k2,…, E_kn] by the text diffusion model.

6. The intelligent question answering method based on diffusion model and large language model according to claim 5, characterized in that: In S3, E k Together with the user question, it is input into the large language model, where E k Insert the attention calculation of the large language model into the corresponding token vector, which includes the following steps: Before generating the answer to Q, the large language model uses the embedding layer in its structure to encode the question Q into a token vector group [e_q1, e_q2, e_q3, …, e_qL], where e_qi is the vector embedding of the i-th token in Q; Add the vector group [E_k1, E_k2, ..., E_kn] with keyword information to the corresponding token vector in the token vector group. Assume n = 1, E_k1 corresponds to the vector representation of the keyword composed of the second and third tokens, then each token vector of e_q2 and e_q3 will be added with E_k1: new_e_q2=e_q2+E_k1; new_e_q3=e_q3+E_k Therefore, the token vector group of the large language model is [e_q1, new_e_q2, new_e_q3, …, e_qL].

7. The intelligent question answering method based on diffusion model and large language model according to claim 6, characterized in that: In S3, the degree of attention paid to each word in the question is calculated, and the answer content is output, including the following steps: The large language model passes the token vector group with keyword information added to it to the corresponding "attention module" in its structure, thereby calculating the degree of attention paid to each word in the question; The vectors corresponding to the keywords in the token vector group incorporate keyword information, thereby enhancing the model's attention to the key parts of question Q and outputting the answer content.