A retrieval-based intelligent question answering method for bidding services
By using BERT, ALBERT and GPT-2 models in the search-based intelligent question-answer system, the existing system's lack of semantic representation capabilities and inflexible answer generation in the bidding service platform are solved, and more accurate and logical answer generation is achieved, improving user experience.
Patent Information
- Application Number
- CN202310605367.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-05-26
AI Technical Summary
The existing search-based intelligent question-and-answer system has problems such as insufficient semantic representation capabilities, poor answer lists, incoherence, and excessive templated answer generation in the bidding service platform.
The BERT model is used to calculate the semantic correlation between the questions and sentences, and the ALBERT model is used to predict the order judgment matrix between the answers in the answer set. The answer set is reordered with the help of hierarchical analysis method, and the GPT-2 model is used to generate the answer text for the question.
It improves the accuracy of answers to complex questions, the generated answers are more logical and inferential, and the user experience is improved, which increases the stickiness between the service platform and users.
Smart Images

Figure CN116521849B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of intelligent question answering, and specifically is a retrieval-type intelligent question answering method for bidding services. Background Art
[0002] With the continuous development of artificial intelligence, the retrieval-based intelligent question-answering system has become one of the important technical assets of the bidding service platform.
[0003] A search-based intelligent question-answering system is a question-answering system based on artificial intelligence technology, which aims to answer questions raised by users. The system mainly relies on search engines and data in knowledge bases to answer questions. This type of question-answering system usually contains two main components: a search engine and an answer generator. The search engine is responsible for finding information related to the user's question in the knowledge base, while the answer generator integrates the searched information and generates the final answer.
[0004] Although the accuracy and efficiency of retrieval-based intelligent question answering systems have been greatly improved, their application in bidding service platforms still has some limitations. First, due to the limitations of semantic representation capabilities, existing research cannot capture the deep semantics of questions and bidding manual document sentences, and therefore cannot accurately answer complex questions; second, existing research obtains answer lists through semantic matching, ignoring the logical order between answer lists, resulting in the answer list received by users being unsmooth and incoherent; third, existing answer generators for bidding services usually use template matching and other technologies to integrate information retrieved from knowledge bases or search engines into the final answer, lacking the ability to integrate question and answer knowledge. This makes the generated answers too templated and lack flexibility. Summary of the invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a retrieval-based intelligent question-answering method for bidding services.
[0006] A search-based intelligent question answering method for bidding services, the method specifically comprising the following steps:
[0007] Step 1: Collect questions and bidding manual documents, and segment the documents at the sentence level to obtain the segmented bidding manual document set D ′ ={D l ′ ,l=1,2,…,L}; for the segmented bidding manual document D ′ All elements in D l ′ The sentences of the sentence are combined to obtain a set S containing P sentences = {S p,p=1,2,…,P}; Sample a subset S from the sentence set ′ ={S t ′ ,t=1,2,…,T}∈S as fine-tuning samples of BERT model parameters;
[0008] Step 2: Use the BERT model to select the N sentences most relevant to the question as the Top-N answer list A i ={A ij ,j=1,2,...,N}, where A ij Denotes the i-th question Q i The jth answer of
[0009] Step 3: Use the ALBERT model to calculate the order judgment matrix between the N answers;
[0010] Step 4: Reorder the N answers using the analytic hierarchy process;
[0011] According to the root method of finding the eigenvalue of the judgment matrix in the hierarchical analysis method, solve the i-th question Q i N answers A i ={A ij ,j=1,2,...,N}. Calculate the order judgment matrix M i The product of the elements of each row in the sequence judgment matrix M is obtained. i The product set of N rows of elements:
[0012]
[0013] in, Denotes the i-th question Q i The order judgment matrix M i The product of the elements of the rth row of ; the sequential judgment matrix M is calculated i The product of the N rows of elements M i The set of Nth roots of :
[0014]
[0015] in, Denotes the i-th question Q i Sequential judgment matrix M i The product of the elements in the rth row of The Nth root of the sequential judgment matrix M i The set of Nth roots of the product of the N rows of elements After normalization, we get the i-th question Q i List A of N answers i The order weight set in Indicates the rth answer A ir The order weight of
[0016] Step 5: Input the reordered N answers into the GPT-2 model to generate the answer text for the question;
[0017] Step 6: Use the answer text as the final feedback response, and attach the Top-N answer list as the basis for the response.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] The present invention uses the BERT model to calculate the semantic relevance between questions and sentences. Since the BERT model has a deep neural network and covers multiple attention heads, it can capture more complex semantic features than traditional language models. At the same time, using the tokenization tool that comes with the BERT model to segment and encode questions and sentences is more convenient and faster than the traditional method of segmenting the text first and then encoding it, thereby improving the efficiency of the experiment;
[0020] The present invention uses the ALBERT model to predict the order judgment matrix between the answers in the answer set. Since the calculated order judgment matrix involves multiple logical inferences, compared with the BERT model, while maintaining high accuracy, ALBERT can greatly reduce the number of parameters and improve the efficiency and scalability of the model. At the same time, since ALBERT uses multi-task learning, it performs better than other models in natural language inference tasks;
[0021] The present invention uses the hierarchical analysis method to compare different answers in the answer set in pairs, and uses a specific mathematical method to calculate the weight of each answer, so as to re-order the answer set, so that the re-ordered answers have stronger logic and inference, and finally improve the accuracy of the final generated answer text;
[0022] The present invention uses the GPT-2 model to learn the potential semantic features and logical relationships between discrete answers, connects the discrete answer fragments into a semantically fluent and content-rich answer text, and finally feeds back the response to the user. Since the GPT-2 model has better language expression and language reasoning capabilities than traditional generation models and can automatically complete text generation tasks, the response is faster and more accurate, thereby improving the user experience and increasing the stickiness of the service platform and users. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flow chart of generating answer text of the present invention;
[0024] Figure 2 The present invention uses the BERT model to calculate the question and answer correlation score graph;
[0025] Figure 3 The present invention uses the ALBERT model to predict the N answers in pairs.
[0026] Figure 4 The present invention uses the GPT-2 model to generate the answer text graph. DETAILED DESCRIPTION
[0027] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] See also Figure 1-Figure 4 As shown, the present application provides a retrieval-based intelligent question answering method for bidding services, which specifically includes the following steps:
[0029] Step 1: Collect questions and bidding manual documents and segment the documents into sentences;
[0030] Step 1.1, collect questions and bidding manual documents;
[0031] Collect questions about bidding and tendering on relevant bidding and tendering service platforms, forums and other platforms, as well as bidding and tendering manual documents containing answers to most questions, including: M questions Q = {Q i ,i=1,2,…,M} and L bidding manual documents D={D l ,l=1,2,…,L};
[0032] In the specific implementation, we obtain the bidding questions and bidding manual documents from the relevant bidding service platforms, forums and other websites, and store them in csv format, so that we can read the documents using python. We filter the documents and questions based on the timestamps of the publications, and finally retain M questions and L bidding manual documents;
[0033] Step 1.2: Segment the document at the sentence level;
[0034] According to the punctuation marks such as ".", ";", "!", "?" and "...", the first document D in the bidding manual document D is taken out respectively. l Perform sentence-level segmentation to obtain the segmented bidding manual document set D′={D′ l ,l=1,2,…,L}. For all elements D′ in the segmented bidding manual document D′ l The sentences of the sentence are combined to obtain a set S containing P sentences = {Sp ,p=1,2,…,P};
[0035] In this embodiment, Python is used to store punctuation marks such as ".", ";", "!", "?" and "..." in a list, and then the document set D is read, and the regular expression and re.split() method of the Python third-party library are used to complete the segmentation of the sentences to obtain the segmented bidding manual document set D'. All the sentences in the segmented bidding manual document set D' are unioned to obtain a set S containing P sentences;
[0036] Sample a subset S′ from the sentence set = {S′ t ,t=1,2,…,T}∈S as the fine-tuning sample of BERT model parameters, annotating each sentence S′ in the subset S′ t With each question Q i The correlation is marked as l′ ti =1, irrelevant label is l′ ti = 0, then the true label is a vector l′ of length M t =[l′ t1 ,l′ t2 ,…,l′ tM ], the annotated subset S′ is used to fine-tune the parameters of the BERT model in step 2;
[0037] In this embodiment, the method for selecting a sentence subset is to use Python to read the sentence set S, store it in the form of a list, use the random.randint() method to generate a random index number for the list, and extract a partial subset S′∈S from the sentence set S for manual labeling.
[0038] Step 2: Use the BERT model to select the N sentences most relevant to the question as the Top-N answer list;
[0039] Step 2.1: Concatenate the question and the sentence as the input of the BERT model to calculate the relevance score between the question and the sentence;
[0040] The i-th question Q i With each sentence S in the sentence set S p Splice to form the i-th question-answer pair set QS i ={(Q i ,S p ),p=1,2,…,P}, where (Q i ,S p ) represents the i-th question-answer pair set QS i The pth question-answer pair in ;
[0041] Construct a BERT model consisting of 12 self-attention layers, using a manually annotated sentence set S ′ Fine-tune the model parameters to obtain a fine-tuned BERT model. i ,S p ) is input into the model and the i-th question Q is output i With the pth sentence S p The correlation score α ip :
[0042] α ip =BERT(Q i ,S p )#(1)
[0043] α i ={α ip ,p=1,2,…,P}#(2)
[0044] where α i Denotes the i-th question Q i With P sentences S = {S p ,p=1,2,…,P};
[0045] In the specific implementation, a multi-layer perceptron network is connected to the output end of the BERT model to convert the output vector of the BERT model into a numerical value. i With the annotation subset S ′ Each sentence S in t ′ Splice to form the i-th training question-answer pair set QS i ′ ={(Q i ,S t ′ ), t=1,2,…,T, input the training question-answer pair Qi,St′ into the BERT model, output the predicted label lti′ of the i-th question Qi and St′, optimize the BERT model using the cross entropy loss function of formula (7) and the back propagation algorithm, and obtain a BERT model with strong learning performance that can be used in this experiment;
[0046]
[0047] Step 2.2: Based on the relevance score, select the most relevant N sentences as the Top-N answer list;
[0048] The i-th question Q i The relevance score α with P sentences S i ={α ip,p=1,2,…,P} sort from large to small, select the first N sentences to form the i-th question Q i List of answers A i ={A ij ,j=1,2,...,N}, where A ij Denotes the i-th question Q i The jth answer of
[0049] In this embodiment, the i-th question Q i The relevance score α with the pth sentence ip The corresponding sentence index number p is saved in a dictionary format, with the sentence index number p as the key of the dictionary. The i-th question Q i The relevance score α with the pth sentence ip As the value of the dictionary, get the i-th question Q i The dictionary of relevance scores to the pth sentence:
[0050] D i ={p:α ip ,p=1,2,…,P}#(8)
[0051] Use Python's sorted() method to sort the i-th question Q i The relevance score dictionary D to the pth sentence i Sort from large to small to get the sorted i-th question Q i The dictionary of relevance scores to the pth sentence:
[0052] D′ i ={p′:α ip′ ,p′=1,2,…,P}#(9)
[0053] Using the for loop, output D′ i The key values of the first N elements constitute the i-th question Q i List of Top-N answer index numbers:
[0054] d i = {p′ j ,j=1,2,…N}#(10)
[0055] According to the index number list d i Take the corresponding sentence from the sentence set S to form the i-th question Q i List of answers A i ={A ij ,j=1,2,...,N};
[0056] Step 3: Calculate the order judgment matrix between N answers based on the ALBERT model;
[0057] Initialize the i-th question Q i A set of N alternative answers i The pairwise sequential judgment matrix in Denotes the i-th question Q i The rth answer A ir and the cth answer A ic The relative order value of
[0058] Using nested loop operations, we start from the i-th question Q i List of answers A i ={A ij ,j=1,2,...,N}, take out the r=1,2,…,Nth answer A ir and the c=1,2,…,Nth answer A ic , if A ir With A ic If they are the same, then their order judgment matrix The value is 1; otherwise, the answer is A ir With answer A ic Concatenate them and input them into the ALBERT model to estimate the logical order of the two answers as the values in the sequential judgment matrix:
[0059]
[0060] According to the analytic hierarchy process, continuous values Discretize. hour, hour, hour, hour, hour, hour, hour, hour, hour,
[0061] The nested loop is to convert the i-th question Q i The N answers are combined in disorder, and only the values above the diagonal of the sequential judgment matrix are calculated. times; the order judgment matrix below the diagonal The value of is equal to its symmetry about the diagonal The reciprocal of
[0062] In the specific implementation, it is necessary to connect a multi-layer perceptron network to the output of the ALBERT model to convert the i-th question Q output by the ALBERT model into i The rth answer Air and the cth answer A ic The relative order value of Map it to a value in [0,1), and then use Python's if statement to answer the i-th question Q i The rth answer A ir and the cth answer A ic The relative order value of Perform threshold judgment and reassignment, and finally write the i-th question Q i With N alternative answer set A i The pairwise sequential judgment matrix M i middle;
[0063] Step 4: Reorder the N answers using the analytic hierarchy process;
[0064] According to the root method of finding the eigenvalue of the judgment matrix in the hierarchical analysis method, solve the i-th question Q i N answers A i ={A ij ,j=1,2,...,N}. Calculate the order judgment matrix M i The product of the elements of each row in the sequence judgment matrix M is obtained. i The product set of N rows of elements:
[0065]
[0066] in, Denotes the i-th question Q i The order judgment matrix M i The product of the elements of the rth row of ; the sequential judgment matrix M is calculated i The product of the N rows of elements M i The set of Nth roots of :
[0067]
[0068] in, Denotes the i-th question Q i Sequential judgment matrix M i The product of the elements in the rth row of The Nth root of the sequential judgment matrix M i The set of Nth roots of the product of the N rows of elements After normalization, we get the i-th question Q i List A of N answers i The order weight set in Indicates the rth answer A ir The order weight of
[0069] In this embodiment, traverse the i-th question Q i The order judgment matrix M i For each row in the , take out each element in each row and multiply them to generate the i-th question Q i The order judgment matrix M i The product list of the elements in the rth row of For equation (5), use Python’s built-in function pow() to solve the i-th question Q i Sequential judgment matrix M i The product of the elements in the rth row of The Nth root of Use the MinMaxScaler() method in the python third-party library sklearn to solve the i-th question Q i List A of N answers i The sequence weight set w i ;
[0070] According to the sequential weight set w i Reorder the N answers to get the i-th question Q i Reordered list of N answers A′ i ={A′ ij}, where A′ ij Denotes the i-th question Q i The jth answer after reordering.
[0071] In the specific implementation, the specific operation method is consistent with step 2.2. i List A of N answers i The index number and weight list w i Use dictionary storage, answer index number j is used as the dictionary key, the i-th question Q i The jth answer A ij The order weight As the value of the dictionary, use Python's sorted() method to sort the i-th question Q i Sort the Nth answer order weight dictionary from large to small, and use the for loop to sort the i-th question Q according to the key value of the sorted order weight dictionary. i A set of N answers i Rearrange and get the i-th question Q i Reordered list of N answers A′ i ;
[0072] Step 5: Input the reordered N answers into the GPT-2 model to generate the answer text for the question;
[0073] The i-th question Qi Reordered list of N answers A′ i Each element A′ in ij Concatenate and add "[SEP]", "TL;DR" and "[BOS]" tokens at the end to represent the input answer text x of the GPT-2 model. i :
[0074] x i =A′ i1 ,A′ i2 ,…,A′ iN ,[SEP],TL;DR,[BOS]#(6)
[0075] After setting the maximum length of the answer text to be generated, the i-th question Q i Input answer text x i Input into the GPT-2 model, and the model outputs the i-th question Q i Final answer text y i ;
[0076] In this example, the GPT-2 model is used to generate the i-th question Q i Final answer text y i The specific process is as follows:
[0077] Step 1: First, you need to install the transformers module of huggingface, and use the import statement to import the tokenization tool AutoTokenizer and the text generation model AutoModelForCausalLM;
[0078] Step 2: Use the AutoTokenizer.from_pretrained() method to load the tokenizer of the GPT-2 model; use the AutoModelForCausalLM.from_pretrained() method to load the GPT-2 model;
[0079] Step 3: The i-th question Q i Input answer text x i Input to the GPT-2 model and output the i-th question Q i Final answer text y i :
[0080] y i =GPT-2(x i )#(11)
[0081] Step 6: Use the answer text as the final feedback response, and attach the Top-N answer list as the basis for the response;
[0082] The i-th question Q i Final answer text y i and its N reordered answer list A′ i Merged into one result, the final answer text represents the summarized answer, and the re-ordered answer list represents the supporting evidence.
[0083] In the specific implementation, as shown in formula (12), the final result is stored in the form of a dictionary, and the i-th question Q i The index number i of the dictionary is used as the key of the dictionary, and the i-th question Q i Final answer text y i and its N reordered answer list A′ i Also store it in a dictionary format to get the i-th question Q i The final result:
[0084] Result i ={i:{Answer text":y i ,"Answer basis":A′ i}}#(12);
[0085] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A search-based intelligent question answering method for bidding services, characterized in that: The method specifically comprises the following steps: Step 1: Collect questions and bidding manual documents, and segment the documents at the sentence level to obtain the segmented bidding manual document set D' = {D' l ,l=1,2,…,L}; for all elements D' in the segmented bidding manual document D' l The sentences of the sentence are combined to obtain a set S containing P sentences = {S p ,p=1,2,…,P}; sample a subset S'={S' t ,t=1,2,…,T}∈S as fine-tuning samples of BERT model parameters; Step 2: Use the BERT model to select the N sentences most relevant to the question as the Top-N answer list A i ={A ij ,j=1,2,...,N}, where A ij Denotes the i-th question Q i The jth answer of Step 3: Use the ALBERT model to calculate the order judgment matrix between the N answers; Step 4: Reorder the N answers using the analytic hierarchy process; According to the root method of finding the eigenvalue of the judgment matrix in the hierarchical analysis method, solve the i-th question Q i N answers A i ={A ij ,j=1,2,...,N}; calculate the order judgment matrix M i The product of the elements of each row in the sequence judgment matrix M is obtained. i The product set of N rows of elements: in, Denotes the i-th question Q i The order judgment matrix M i The product of the elements of the rth row of ; the sequential judgment matrix M is calculated i The product of the N rows of elements M i The set of Nth roots of : in, Denotes the i-th question Q i Sequential judgment matrix M i The product of the elements in the rth row of The Nth root of the sequential judgment matrix M i The set of Nth roots of the product of the N rows of elements After normalization, we get the i-th question Q i List A of N answers i The sequence weight set in Indicates the rth answer A ir The order weight of Step 5: Input the reordered N answers into the GPT-2 model to generate the answer text for the question; Step 6: Use the answer text as the final feedback response, and attach the Top-N answer list as the basis for the response.
2. A search-based intelligent question-answering method for bidding services according to claim 1, characterized in that: The specific method of sentence-level segmentation of documents is as follows: S11, according to the punctuation marks ".", ";", "!", "?" and "...", respectively take out the first document D in the bidding manual document D l Perform sentence-level segmentation to obtain the segmented bidding manual document set D' = {D' l ,l=1,2,…,L}; for all elements D' in the segmented bidding manual document D' l The sentences of the sentence are combined to obtain a set S containing P sentences = {S p ,p=1,2,…,P}; S12, sampling a subset S'={S' t ,t=1,2,…,T}∈S as a fine-tuning sample of the BERT model parameters, annotating each sentence S' in the subset S' t With each question Q i The correlation is marked as l' ti =1, irrelevant mark is l' ti = 0, then the true label is a vector l' of length M t =[l' t1 ,l' t2 ,…,l' tM ], the annotated subset S' is used to fine-tune the parameters of the BERT model in step 2.
3. A search-based intelligent question-answering method for bidding services according to claim 1, characterized in that: The specific way to select the answer list using the BERT model is as follows: Step 2.1: Concatenate the question and the sentence as the input of the BERT model to calculate the relevance score between the question and the sentence; The i-th question Q i With each sentence S in the sentence set S p Splice to form the i-th question-answer pair set QS i ={(Q i ,S p ),p=1,2,…,P}, where (Q i ,S p ) represents the i-th question-answer pair set QS i The pth question-answer pair in ; A BERT model consisting of 12 self-attention layers is constructed, and the model parameters are fine-tuned using the manually annotated sentence set S' to obtain a fine-tuned BERT model. i ,S p ) is input into the model and the i-th question Q is output i With the pth sentence S p The correlation score α ip : α ip =BERT(Q i ,S p )#(1) a i ={a ip ,p=1,2,…,P}#(2) where α i Denotes the i-th question Q i With P sentences S = {S p ,p=1,2,…,P}; Step 2.2: Based on the relevance score, select the most relevant N sentences as the Top-N answer list; The i-th question Q i The relevance score α with P sentences S i ={α ip ,p=1,2,…,P} sort from large to small, select the first N sentences to form the i-th question Q i List of answers A i ={A ij ,j=1,2,...,N}, where A ij Denotes the i-th question Q i The jth answer of .
4. A search-based intelligent question-answering method for bidding services according to claim 1, characterized in that: The specific method of using the ALBERT model to calculate the order judgment matrix between N answers is as follows: Initialize the i-th question Q i A set of N alternative answers i The pairwise sequential judgment matrix in Denotes the i-th question Q i The rth answer A ir and the cth answer A ic The relative order value of Using nested loop operations, we start from the i-th question Q i List of answers A i ={A ij ,j=1,2,...,N}, take out the r=1,2,…,Nth answer A ir and the c=1,2,…,Nth answer A ic , if A ir With A ic If they are the same, then their order judgment matrix The value is 1; otherwise, the answer is A ir With answer A ic Concatenate them and input them into the ALBERT model to estimate the logical order of the two answers as the values in the sequential judgment matrix: According to the analytic hierarchy process, continuous values Discretize; specify hour, hour, hour, hour, hour, hour, hour, hour, hour, The nested loop is to convert the i-th question Q i The N answers are combined in disorder, and only the values above the diagonal of the sequential judgment matrix are calculated. times; the order judgment matrix below the diagonal The value of is equal to its symmetry about the diagonal The reciprocal of .
5. A search-based intelligent question-answering method for bidding services according to claim 1, characterized in that: The specific method of generating the answer text to the question is: The i-th question Q i Reordered list of N answers A' i Each element A' ij Concatenate and add "[SEP]", "TL;DR" and "[BOS]" tokens at the end to represent the input answer text x of the GPT-2 model. i : x i =A′ i1 ,A′ i2 ,…,A′ iN ,[SEP],TL;DR,[BOS]#(6) After setting the maximum length of the answer text to be generated, the i-th question Q i Input answer text x i Input into the GPT-2 model, and the model outputs the i-th question Q i Final answer text y i .
6. A search-based intelligent question answering method for bidding services according to claim 1, characterized in that: Step 6 is specifically as follows: The i-th question Q i Final answer text y i and its N reordered answer list A' i Merged into one result, the final answer text represents the summarized answer, and the re-ordered answer list represents the supporting evidence.
Citation Information
Patent Citations
FAQ question similarity calculation method and system
CN111581354A
Retrieval type intelligent question and answer system and method for coal mine safety regulations
CN114020862A