A real-time speech stream and text dialogue interaction system based on large language models
By combining voice input, time series analysis, perceptual marking and confidence analysis modules, answers are preferred in the corpus, and when they are not found, they call large language models to generate responses, which solves the problem of difficult to quickly answer professional problems in the existing technology, and realizes an efficient, intelligent and personalized dialogue system.
Patent Information
- Application Number
- CN202411571740.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-06
AI Technical Summary
The existing real-time speech streaming and text dialogue interaction systems of large language models are difficult to combine their own designed corpus and large language models, making it difficult for the generated answers to quickly respond to professional questions in a targeted and fast manner.
The speech input module, time series analysis module, perception marking module, answer generation module and confidence analysis module are adopted. Through NLP technology and deep learning model, combined with user attributes, static logs and key fields of the question, answers are preferred in the corpus. When it is not found, a large language model is called to generate responses, and the accuracy of the answer is ensured through the confidence threshold mechanism.
It improves the system's ability to respond quickly to professional questions, reduces response time, enhances the ability to adapt to complex or new questions, provides personalized answers that are more in line with user needs, and improves interaction quality and system processing efficiency.
Smart Images

Figure CN119539076B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of dialogue interaction systems, and specifically relates to a real-time speech stream and text dialogue interaction system based on a large language model. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the progress of large language models such as GPT series and BERT, natural language processing (NLP) has been widely applied in multiple fields. Real-time speech stream and text dialogue interaction systems are the products of this trend. They combine speech recognition, natural language understanding, and generation technologies to provide users with a more intuitive and efficient interaction method. A real-time speech stream and text dialogue interaction system based on a large language model can receive the user's speech input, convert it into text, analyze the user's intention, generate corresponding answers according to the context, and support real-time feedback.
[0003] Most of the existing real-time speech stream and text dialogue interaction systems based on large language models directly generate responses through the large language model, and it is difficult to combine the corpus designed by themselves with the large language model, making it difficult for the generated answers to quickly respond to professional questions in a targeted manner. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a real-time speech stream and text dialogue interaction system based on a large language model, which is used to solve the technical problem that it is difficult to combine the corpus designed by itself with the large language model, making it difficult for the generated answers to quickly respond to professional questions in a targeted manner.
[0005] To solve the above problems, the first aspect of the present invention provides a real-time speech stream and text dialogue interaction system based on a large language model, including:
[0006] A speech input module: used to receive the speech or text information input by the client, and convert the speech input by the user into corresponding text information;
[0007] A time series analysis module: through time series analysis, perform historical tracking on user attributes, user static logs, and user questions;
[0008] A perception marking module: extract key fields of user attributes, user static logs, and user questions tracked by the time series analysis module through NLP technology, and set user tags according to the key fields of user attributes, user static logs, and user questions;
[0009] Answer generation module: It is used to query answers in the corpus according to the user tags of the user and the user question data currently input by the user. For questions with answers found in the corpus, response data is generated based on the data in the corpus; for questions without answers found in the corpus, a large language model is called to generate response data;
[0010] Confidence analysis module: It is used to save the context information in the current conversation through a deep learning model, generate semantic features of the context information, and analyze the confidence of the answers generated by the answer generation module. Through a confidence threshold mechanism, responses that meet the confidence requirements are output;
[0011] Output module: It converts the finally generated answer into natural language and synthesizes it to return to the user.
[0012] As a further solution of the present invention: The time series analysis module performs historical tracking on user attributes, user static logs, and user questions through time series analysis, including the following steps:
[0013] Obtain user attribute and user static log data. User attribute data includes: age, gender, and major. User static log data includes: the interaction content of each conversation with the system, and the satisfaction feedback of the user on the system's answers;
[0014] Clean the obtained user attribute and user static log data to remove invalid data; such as duplicate records and missing values;
[0015] Standardize data from different sources and perform unified format processing;
[0016] Extract the time series of user conversations through time series analysis technology, and determine user conversations with a frequency of greater than or equal to 8 times per hour as conversations belonging to the same cycle;
[0017] Screen the conversations of the last five cycles and the corresponding user static log data as the user static log data for tracking and detection.
[0018] As a further solution of the present invention: The perception marking module extracts key fields of user attributes, user static logs, and user questions tracked by the time series analysis module through NLP technology, and sets user tags according to the key fields of user attributes, user static logs, and user questions, including the following steps:
[0019] Extract key fields of user attributes, user static logs, and user questions through the spaCy library of NLP technology;
[0020] Based on the extracted user attributes, user static logs, and keyword fields of the user question, use the keyword fields of the user attributes and user static logs as user tags;
[0021] According to the filtered conversations in the most recent five cycles, construct a list of keyword fields corresponding to the user question, set corresponding key-value pairs for the list of keyword fields of the user question in each cycle, and add the key-value pairs corresponding to the list to the user tags.
[0022] As a further solution of the present invention: the answer generation module queries for answers in the corpus according to the user tags of the user and the user question data currently input by the user. For questions for which answers are found in the corpus, generate response data according to the data in the corpus, including the following steps:
[0023] Obtain the text data of the latest professional books, academic papers, and academic forums, set professional knowledge Q&A pairs according to the collected text data, and extract the keyword fields of each Q&A pair;
[0024] Extract keyword fields from the user's question;
[0025] Search for relevant questions in the professional knowledge Q&A pairs in the corpus according to the keyword fields in the user question;
[0026] Convert the keyword fields in the user question and the keyword fields of the Q&A pairs in the corpus into vector representations through the Word2Vec model;
[0027] Calculate the semantic matching degree between the user question and the Q&A pairs in the corpus according to the cosine similarity of the word vectors between the keyword fields in the user question and the keyword fields of the Q&A pairs in the corpus;
[0028] Among the Q&A pairs where the semantic matching degree between the user question and the Q&A pairs in the corpus is greater than the preset threshold, filter out the Q&A pair with the maximum semantic matching degree;
[0029] Generate response data according to the answer data corresponding to the Q&A pair with the maximum semantic matching degree.
[0030] As a further solution of the present invention: the semantic matching degree between the user question and the Q&A pairs in the corpus is calculated through the following formula:
[0031]
[0032] where Ex is the semantic matching degree between the user question and the Q&A pairs in the corpus, and COSX i is the average value of the cosine similarity of the word vectors between the i-th keyword field in the user question and the keyword field of the Q&A pair in the corpus, i ∈ (1, 2,..., n), and n is the total number of keyword fields in the user question;
[0033] As a further solution of the present invention: The confidence analysis module includes:
[0034] A keyword generation unit: It is used to analyze and generate keyword vectors in the generated response data through the LSTM layer and Transformer layer constructed in the deep learning model, and account for the attention weights of the keyword vectors of each question and answer in the keyword vector sequence;
[0035] An analysis unit: It is used to analyze the confidence of the response data generated by the answer generation module according to the obtained attention weights, and output a response that meets the confidence requirements through a confidence threshold mechanism;
[0036] Among them, the confidence threshold mechanism includes: According to the confidence threshold of the response data, when it is detected that the confidence of the response data is lower than the confidence threshold, the user's dialogue intention is actively confirmed, otherwise, the intention is not confirmed.
[0037] As a further solution of the present invention: The keyword generation unit analyzes and generates keyword vectors in the generated response data through the LSTM layer and Transformer layer constructed in the deep learning model, and accounts for the attention weights of the keyword vectors of each question and answer in the keyword vector sequence, including the following steps:
[0038] Extract keywords from the context information of the current dialogue and convert the keywords into vector representations;
[0039] Construct an LSTM layer, with the input being the keyword vectors of the context information. The LSTM layer processes the time when the keyword vectors of the current dialogue appear in the context through its internal memory unit and gating mechanism, generates a time series of keyword vectors, and adds time series labels to each keyword according to the time series of keyword vectors;
[0040] The output of the LSTM layer is a sequence of keyword vectors with time series labels;
[0041] Construct a Transformer layer. The input of the Transformer layer is the output of the LSTM layer. The Transformer layer analyzes the keyword vectors in the latest generated response data of the current dialogue in the keyword vector sequence through the Self-Attention self-attention mechanism, and accounts for the attention weights of the keyword vector sequence;
[0042] The Transformer layer outputs the keyword vectors in the latest generated response data of the current dialogue, accounting for the attention weights of the keyword vector sequence;
[0043] The output of the Transformer layer is input into the fully connected layer of the deep learning model to obtain the keyword vectors in the newly generated response data, which account for the attention weights of the keyword vectors of each question and answer in the keyword vector sequence.
[0044] As a further solution of the present invention: the analysis unit analyzes the confidence of the response data generated by the answer generation module according to the obtained attention weights, including the following steps:
[0045] According to the keyword vector sequence with time series tags added, analyze the cosine similarity between the keyword vectors in the response data and the word vectors in the keyword vector sequence, and calculate the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence;
[0046] According to the obtained semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence and the attention weights generated by the keyword generation unit, analyze the confidence of the response data generated by the answer generation module.
[0047] As a further solution of the present invention: calculate the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence through the following formula:
[0048]
[0049] where, Ey j is the semantic matching degree between the j-th keyword in the response data and the word vectors in the keyword vector sequence, and COSy j is the average value of the cosine similarity between the j-th keyword vector in the response data and the word vectors in the keyword vector sequence;
[0050] As a further solution of the present invention: analyze the confidence of the response data generated by the answer generation module through the following formula:
[0051]
[0052] where, C is the confidence of the response data generated by the answer generation module, and K j is the average value of the attention weights of the j-th keyword vector in the response data accounting for the keyword vectors of each question and answer in the keyword vector sequence, j ∈ (1, 2,..., m), and m is the total number of keywords in the response data.
[0053] Compared with the prior art, the beneficial effects of the present invention are:
[0054] The present invention reduces the response time by quickly returning predefined answers for questions that can find answers in the corpus. By preferentially using the existing knowledge base, the burden on the large language model can be reduced, and the overall processing efficiency of the system can be improved. The information obtained from the corpus is usually more professional data, which can improve the accuracy and reliability of the answers. At the same time, by combining user tags to provide information more in line with their needs or background, the responses can be made more in line with user expectations. When answers cannot be found in the corpus, the large language model can be called to generate responses, enhancing the system's adaptability to complex or new questions.
[0055] The present invention combines the advantages of deep learning technology in context understanding, semantic feature extraction, and confidence evaluation, and uses a confidence threshold mechanism to ensure that the output content is both accurate and meets user needs. Ultimately, an efficient, intelligent, and user-centered dialogue system is achieved, which not only improves the interaction quality but also enhances the system's ability to handle diverse scenarios and complex problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0057] Figure 1 It is a schematic diagram of the system framework of the present invention;
[0058] Figure 2 It is a schematic diagram of the framework of the confidence analysis module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0060] Please refer to Figure 1 - Figure 2 , the first aspect embodiment of the present invention provides a real-time voice stream and text dialogue interaction system based on a large language model, including:
[0061] A voice input module: used to receive voice or text information input by the client and convert the voice input by the user into corresponding text information;
[0062] Time Series Analysis Module: Through time series analysis, historical tracking of user attributes, user static logs, and user questions is performed;
[0063] Perception Tagging Module: Through NLP technology, extract key fields of user attributes, user static logs, and user questions tracked by the time series analysis module, and set user tags based on the key fields of user attributes, user static logs, and user questions;
[0064] Answer Generation Module: Used to query answers in the corpus according to the user's user tags and the user question data currently input by the user. For questions with answers found in the corpus, generate response data based on the data in the corpus; for questions without answers found in the corpus, call the large language model to generate response data;
[0065] Confidence Analysis Module: Used to save the context information in the current conversation through a deep learning model, generate semantic features of the context information, and analyze the confidence of the answers generated by the answer generation module. Through the confidence threshold mechanism, output responses that meet the confidence requirements;
[0066] Output Module: Convert the finally generated answer into natural language and synthesize it to return to the user.
[0067] Specifically, in this embodiment, the set perception tagging module extracts key fields of user attributes, user static logs, and user questions tracked by the time series analysis module, and sets user tags based on the key fields of user attributes, user static logs, and user questions;
[0068] By analyzing user attributes and user static logs, it is convenient to provide more personalized recommendations or answers that meet their needs, improving user satisfaction. Through key field extraction, the current problems and needs of users can be quickly understood, thus accelerating the speed of providing solutions. Combining time series data can better understand the focus of the user's conversation within a specific time period.
[0069] In this embodiment, the answer generation module queries answers in the corpus according to the user's user tags and the user question data currently input by the user. For questions with answers found in the corpus, generate response data based on the data in the corpus; for questions without answers found in the corpus, call the large language model to generate response data.
[0070] For questions with answers that can be found in the corpus, quickly return the predefined answers, thereby reducing the response time. By preferentially using the existing knowledge base, the burden on the large language model can be reduced, improving the overall processing efficiency of the system.
[0071] The information obtained from the corpus is usually more professional data, which can improve the accuracy and reliability of the answers.
[0072] At the same time, combined with user tags, it provides information that better meets their needs or background, making the response more in line with user expectations.
[0073] When an answer cannot be found in the corpus, it can call a large language model to generate a response, enhancing the system's adaptability to complex or new questions. The large language model called can be selected from the existing authorized large language models.
[0074] This hybrid approach not only improves the system response speed and accuracy but also allows for flexible handling of various types of questions. Combining the expertise from the corpus and the dynamic generation ability from the large language model can provide users with more comprehensive, personalized, and efficient services. This strategy can ensure the consistency of information while adapting to changing needs.
[0075] In this embodiment, the confidence analysis module saves the context information in the current conversation through a deep learning model, generates semantic features of the context information, and analyzes the confidence of the answer generated by the answer generation module. Through the confidence threshold mechanism, it outputs a response that meets the confidence requirements;
[0076] By saving the conversation context information, it can help the model better understand the user's intentions and needs, thus providing more relevant answers. By maintaining the context, it can effectively avoid misunderstandings or incorrect responses caused by the lack of background information. The deep learning model can capture complex semantic relationships, analyze the confidence of the generated answer, facilitate the screening of answers with qualified confidence, and improve the overall response quality of the system.
[0077] Integrating the advantages of deep learning technology in context understanding, semantic feature extraction, and confidence evaluation, through the confidence threshold mechanism to ensure that the output content is both accurate and meets user needs. Finally, an efficient, intelligent, and user-centered dialogue system is realized, which not only improves the interaction quality but also enhances the system's ability to handle diverse scenarios and complex problems.
[0078] In one embodiment of the present invention, the time series analysis module performs historical tracking on user attributes, user static logs, and user questions through time series analysis, including the following steps:
[0079] Obtain user attribute and user static log data. User attribute data includes: age, gender, and major. User static log data includes: the interaction content with the system in each conversation, and the satisfaction feedback of the user on the system's answer;
[0080] Clean the acquired user attributes and user static log data to remove invalid data, such as duplicate records and missing values;
[0081] Standardize data from different sources and process them into a unified format;
[0082] Extract the time series of user conversations through time series analysis technology, and determine that user conversations with a frequency greater than or equal to 8 times per hour belong to conversations in the same period;
[0083] Screen the conversations and corresponding user static log data of the last five periods as the user static log data for tracking and detection.
[0084] In one embodiment of the present invention, the perception marking module extracts the keyword fields of user attributes, user static logs, and user questions tracked by the time series analysis module through NLP technology, and sets user tags according to the keyword fields of user attributes, user static logs, and user questions, including the following steps:
[0085] Extract the keyword fields of user attributes, user static logs, and user questions through the spaCy library of NLP technology;
[0086] According to the extracted keyword fields of user attributes, user static logs, and user questions, use the keyword fields of user attributes and user static logs as user tags;
[0087] According to the screened conversations of the last five periods, construct a keyword field list for the corresponding user questions, set corresponding key-value pairs for the keyword field list of user questions in each period, and add the key-value pairs corresponding to the list to the user tags.
[0088] In one embodiment of the present invention, the answer generation module queries for answers in the corpus according to the user tags of the user and the user question data currently input by the user. For questions for which answers are found in the corpus, generate response data according to the data in the corpus, including the following steps:
[0089] Obtain the text data of the latest professional books, academic papers, and academic forums, set professional knowledge Q&A pairs according to the collected text data, and extract the keyword fields of each Q&A pair;
[0090] Extract the keyword fields from the user's question;
[0091] Search for relevant questions in the professional knowledge Q&A pairs in the corpus according to the keyword fields in the user's question;
[0092] Convert the keyword fields in the user's question and the keyword fields of the Q&A pairs in the corpus into vector representations through the Word2Vec model;
[0093] Calculate the semantic matching degree between the user question and the Q&A pairs in the corpus based on the cosine similarity of the word vectors between the keyword fields in the user question and the keyword fields of the Q&A pairs in the corpus;
[0094] Among the Q&A pairs with the semantic matching degree between the user question and the Q&A pairs in the corpus greater than the preset threshold, select the Q&A pair with the maximum semantic matching degree;
[0095] Generate response data according to the response data corresponding to the Q&A pair with the maximum semantic matching degree.
[0096] In one embodiment of the present invention, the semantic matching degree between the user question and the Q&A pairs in the corpus is calculated through the following formula:
[0097]
[0098] where Ex is the semantic matching degree between the user question and the Q&A pairs in the corpus, and COSX i is the average value of the cosine similarities of the word vectors between the i-th keyword field in the user question and the keyword fields of the Q&A pairs in the corpus, i ∈ (1, 2,..., n), and n is the total number of keyword fields in the user question;
[0099] In one embodiment of the present invention, the confidence analysis module includes:
[0100] Keyword generation unit: used to analyze and generate the attention weights of the keyword vectors in the generated response data to the keyword vectors of each Q&A in the keyword vector sequence through the LSTM layer and Transformer layer constructed in the deep learning model;
[0101] Analysis unit: used to analyze the confidence of the response data generated by the answer generation module according to the obtained attention weights, and output a response that meets the confidence requirement through the confidence threshold mechanism;
[0102] Among them, the confidence threshold mechanism includes: according to the confidence threshold of the response data, when it is detected that the confidence of the response data is lower than the confidence threshold, actively confirm the conversation intention with the user, otherwise, do not confirm the intention.
[0103] Specifically, in this embodiment, it is set according to the intermediate value of the average value and the minimum value of the confidence of the response data in which the user has not shown understanding defects in the user's historical conversation. If there is no historical conversation, or the historical conversation is less than 5 times, the confidence threshold of the initial response data is set to 0.5, and the response can be preliminarily screened before being returned to the user.
[0104] In one embodiment of the present invention, the keyword generation unit analyzes the keyword vectors in the generated response data through the LSTM layer and the Transformer layer constructed in the deep learning model, and obtains the attention weights of the keyword vectors of each question and answer in the keyword vector sequence, including the following steps:
[0105] Extract keywords from the context information of the current conversation and convert the keywords into vector representations;
[0106] Construct an LSTM layer with the keyword vectors of the context information as the input. The LSTM layer processes the time when the keyword vectors of the current conversation appear in the context through its internal memory units and gating mechanisms, generates a time series of keyword vectors, and adds time series labels to each keyword according to the time series of keyword vectors;
[0107] The output of the LSTM layer is a sequence of keyword vectors with time series labels added;
[0108] Construct a Transformer layer with the output of the LSTM layer as the input. The Transformer layer analyzes the keyword vectors in the latest generated response data of the current conversation in the keyword vector sequence through the Self-Attention mechanism and obtains the attention weights of the keyword vectors in the keyword vector sequence;
[0109] The Transformer layer outputs the attention weights of the keyword vectors in the latest generated response data of the current conversation in the keyword vector sequence;
[0110] The output of the Transformer layer is input into the fully connected layer of the deep learning model to obtain the attention weights of the keyword vectors in the latest generated response data for each question and answer in the keyword vector sequence.
[0111] In one embodiment of the present invention, the analysis unit analyzes the confidence of the response data generated by the answer generation module according to the obtained attention weights, including the following steps:
[0112] According to the sequence of keyword vectors with time series labels added, analyze the cosine similarity between the keyword vectors in the response data and the word vectors in the keyword vector sequence, and calculate the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence;
[0113] According to the obtained semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence and the attention weights generated by the keyword generation unit, analyze the confidence of the response data generated by the answer generation module.
[0114] In one embodiment of the present invention, the semantic matching degree between the keywords in the response data and the word vectors of the keyword vector sequence is calculated by the following formula:
[0115]
[0116] where Ey j is the semantic matching degree between the j-th keyword in the response data and the word vectors of the keyword vector sequence, and COSy j is the average value of the cosine similarity between the j-th keyword vector in the response data and the word vectors of the keyword vector sequence;
[0117] In one embodiment of the present invention, the confidence of the response data generated by the answer generation module is analyzed by the following formula:
[0118]
[0119] where C is the confidence of the response data generated by the answer generation module, and K j is the mean of the attention weights of the j-th keyword vector in the response data to the keyword vectors of each Q&A in the keyword vector sequence, j ∈ (1, 2,..., m), and m is the total number of keywords in the response data.
[0120] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A real-time speech stream and text dialogue interaction system based on a large language model, characterized in that, Including: A voice input module: used to receive voice or text information input by the client and convert the voice input by the user into corresponding text information; A time series analysis module: through time series analysis, perform historical tracking on user attributes, user static logs, and user questions; A perception marking module: extract key fields of user attributes, user static logs, and user questions tracked by the time series analysis module through NLP technology, and set user tags according to the key fields of user attributes, user static logs, and user questions; An answer generation module: used to query answers in the corpus according to the user's user tags and the user question data currently input by the user. For questions where answers are found in the corpus, generate response data according to the data in the corpus; For questions where answers cannot be found in the corpus, call a large language model to generate response data; A confidence analysis module: used to save the context information in the current conversation through a deep learning model, generate semantic features of the context information, and analyze the confidence of the answers generated by the answer generation module. Through a confidence threshold mechanism, output responses that meet the confidence requirements; An output module: convert the finally generated answer into natural language and synthesize it to return to the user; Among them, the keyword generation unit of the confidence analysis module analyzes the keyword vectors in the generated response data through the LSTM layer and Transformer layer built in the deep learning model, and accounts for the attention weights of the keyword vectors of each question and answer in the keyword vector sequence, including: Extract keywords from the context information of the current conversation and convert the keywords into vector representations; Build an LSTM layer, with the input being the keyword vectors of the context information. The LSTM layer processes the time when the keyword vectors of the current conversation appear in the context through its internal memory unit and gating mechanism, generates a time series of keyword vectors, and adds time series tags to each keyword according to the time series of keyword vectors; The output of the LSTM layer is a sequence of keyword vectors with time series tags added; Build a Transformer layer, with the input of the Transformer layer being the output of the LSTM layer. The Transformer layer analyzes the keyword vectors in the latest generated response data of the current conversation in the keyword vector sequence through the Self-Attention self-attention mechanism and accounts for the attention weights of the keyword vector sequence; The Transformer layer outputs the keyword vectors in the latest generated response data of the current conversation, accounting for the attention weights of the keyword vector sequence; The output of the Transformer layer is input into the fully connected layer of the deep learning model to obtain the keyword vectors in the latest generated response data, accounting for the attention weights of the keyword vectors of each question and answer in the keyword vector sequence; The analysis unit of the confidence analysis module analyzes the confidence of the response data generated by the answer generation module according to the obtained attention weights, including: According to the keyword vector sequence with added time series tags, analyze the cosine similarity between the keyword vectors in the response data and the word vectors in the keyword vector sequence, and calculate the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence; among them, the calculation of the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence is obtained through the average value of the cosine similarity between the keyword vectors and the word vectors in the keyword vector sequence. According to the obtained semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence and the attention weights generated by the keyword generation unit, analyze the confidence of the response data generated by the answer generation module; among them, analyzing the confidence of the response data generated by the answer generation module is obtained through weighted averaging of the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence, and the weights are set according to the attention weights generated by the keyword generation unit.
2. The real-time speech stream and text dialogue interaction system based on a large language model according to claim 1, wherein, The time series analysis module performs historical tracking on user attributes, user static logs, and user questions through time series analysis, including the following steps: Obtain user attribute and user static log data. User attribute data includes: age, gender, and major. User static log data includes: the interaction content of each conversation with the system, and the satisfaction feedback of the user on the system's answer. Clean the obtained user attribute and user static log data to remove invalid data. Invalid data includes: duplicate records and missing values. Standardize the data from different sources and perform unified format processing. Extract the time series of user conversations through time series analysis technology, and determine that user conversations with a frequency greater than or equal to 8 times per hour belong to conversations in the same period. Screen the conversations of the last five periods and the corresponding user static log data as the user static log data for tracking detection.
3. A real-time speech stream and text dialogue interaction system based on a large language model according to claim 2, characterized in that, The perception marking module extracts the key fields of user attributes, user static logs, and user questions tracked by the time series analysis module through NLP technology, and sets user tags according to the key fields of user attributes, user static logs, and user questions, including the following steps: Extract the key fields of user attributes, user static logs, and user questions through the spaCy library of NLP technology. According to the extracted key fields of user attributes, user static logs, and user questions, use the key fields of user attributes and user static logs as user tags. According to the screened conversations of the last five periods, construct a list of key fields for the corresponding user questions, set corresponding key-value pairs for the list of key fields of user questions in each period, and add the corresponding key-value pairs of the list to the user tags.
4. A real-time speech stream and text dialogue interaction system based on a large language model according to claim 1, characterized in that, The answer generation module queries for answers in the corpus according to the user's user tags and the user question data currently input by the user. For questions where answers are found in the corpus, generate response data according to the data in the corpus, including the following steps: Obtain the text data of the latest professional books, academic papers and academic forums, set professional knowledge Q&A pairs according to the collected text data, and extract the key fields of each Q&A pair; Extract the key fields from the user's question; Search for relevant questions in the professional knowledge Q&A pairs in the corpus according to the key fields in the user's question; Convert the key fields in the user's question and the key fields of the Q&A pairs in the corpus into vector representations through the Word2Vec model; Calculate the semantic matching degree between the user's question and the Q&A pairs in the corpus according to the cosine similarity of the word vectors between the key fields in the user's question and the key fields of the Q&A pairs in the corpus; Among the Q&A pairs with the semantic matching degree between the user's question and the Q&A pairs in the corpus greater than the preset threshold, screen out the Q&A pair with the maximum semantic matching degree; Generate response data according to the response data corresponding to the Q&A pair with the maximum semantic matching degree.
5. The real-time speech stream and text dialogue interaction system based on a large language model according to claim 4, wherein, The semantic matching degree between the user's question and the Q&A pairs in the corpus is calculated through the following formula: Among them, Ex is the semantic matching degree between the user's question and the Q&A pairs in the corpus, and COSX i is the average value of the cosine similarity of the word vectors between the i-th keyword field in the user's question and the keyword fields of the Q&A pairs in the corpus, where i ∈ (1, 2, …, n), and n is the total number of keyword fields in the user's question.
6. The real-time speech stream and text conversation interaction system based on a large language model according to claim 1, characterized in that, The confidence analysis module includes: The keyword generation unit: used to analyze and generate the attention weights of the keyword vectors in the generated response data in the keyword vector sequence of each Q&A through the LSTM layer and Transformer layer constructed in the deep learning model; The analysis unit: used to analyze the confidence of the response data generated by the answer generation module according to the obtained attention weights, and output the response that meets the confidence requirement through the confidence threshold mechanism; Among them, the confidence threshold mechanism includes: according to the confidence threshold of the response data, when it is detected that the confidence of the response data is lower than the confidence threshold, actively confirm the conversation intention with the user, otherwise, do not confirm the intention.
7. A real-time speech stream and text dialogue interaction system based on a large language model according to claim 1, characterized in that, Calculate the semantic matching degree between the keywords in the response data and the word vectors in the keyword vector sequence through the following formula: Among them, Ey j is the semantic matching degree between the j-th keyword in the response data and the word vectors in the keyword vector sequence, and COSy j is the average value of the cosine similarity between the j-th keyword vector in the response data and the word vectors in the keyword vector sequence.
8. The real-time speech stream and text dialogue interaction system based on a large language model according to claim 7, wherein, Analyze the confidence of the response data generated by the answer generation module through the following formula: Among them, C is the confidence of the response data generated by the answer generation module, and K j is the mean of the attention weights of the j-th keyword vector in the response data to the keyword vectors of each Q&A in the keyword vector sequence, where j ∈ (1, 2, …, m), and m is the total number of keywords in the response data.
Citation Information
Patent Citations
Machine reading method and device based on transformer and lstm and readable storage medium
CN110866098A
Customer service dialogue method based on artificial intelligence
CN111797202A
Local dialect voice intelligent identification and question answering method, system and device and medium
CN117198267A