Intelligent question answering system and method based on large model
By adopting large models and emotion recognition technology in the intelligent question-and-answer system, the problem of traditional systems being difficult to understand natural language and neglecting user emotions is solved, and more accurate and personalized responses are achieved, improving the user experience.
Patent Information
- Application Number
- CN202510161056.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional intelligent question-and-answer systems are difficult to accurately understand the diversity and complexity of natural language, and due to the scale of the knowledge base, the reply content may be one-sided, outdated or inaccurate, ignoring the user's emotional state and psychological needs, resulting in poor user experience.
The intelligent question-and-answer system based on the big model is adopted, and the user's question description data collection module, the reply text generation module, the user's question description data collection module and the user's question emotional recognition module are used to obtain the question description and question text entered by the user, and the big model is used to generate the reply text, and the user's emotional state is judged through the emotion recognition module. If you are not satisfied, a pop-up window will be generated to prompt whether manual reply is required.
It improves the accuracy of question comprehension and relevance and timeliness of reply content, provides personalized and humanized service experience through emotional perception, and improves user experience and service quality.
Smart Images

Figure CN120030127A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent question answering, and more specifically, to an intelligent question answering system and method based on a large model. Background Art
[0002] In today's era of digital information explosion, people's demand for fast and accurate information is growing. As an efficient information interaction tool, the intelligent question-answering system has been widely used in many fields such as education, medical care, customer service, etc. due to its immediacy and accuracy, significantly improving the efficiency of information acquisition and user experience.
[0003] Traditional systems rely on pre-set rules and limited knowledge bases, and it is difficult to accurately understand the diversity and complexity of natural language. In addition, due to the size of the knowledge base, the responses of traditional systems may be one-sided, outdated or inaccurate. In professional fields such as law, medicine, literature, and teaching materials, it is difficult to provide comprehensive and up-to-date professional knowledge. In addition, traditional systems usually focus on the mechanical matching of questions and answers, ignoring the emotional state and psychological needs of users. This one-way information transmission method may lead to a poor user experience, especially when users express dissatisfaction or confusion and fail to respond in time.
[0004] Therefore, an intelligent question-answering solution based on a large model is desired. Summary of the invention
[0005] In order to solve the above technical problems, this application is proposed. The embodiments of this application provide an intelligent question answering system and method based on a large model.
[0006] According to one aspect of the present application, a large model-based intelligent question-answering system is provided, which includes: a user question description data collection module, which is used to obtain a question description input by a user; a reply text generation module, which is used to input the question description into a dialogue engine based on a large model to obtain a reply text; a user re-question description data collection module, which is used to obtain a re-question text description input by the user for the reply text; a user re-question emotion recognition module, which is used to perform emotion recognition on the re-question text description to obtain an emotion recognition result; a pop-up window triggering module, which is used to generate a pop-up window prompt whether to manually reply in response to the emotion recognition result being unsatisfactory; The user then asks the emotion recognition module questions, including: A text description sentence encoding unit, used for performing sentence processing and semantic encoding on the re-question text description to obtain a sequence distribution of semantic encoding vectors of the re-question sentence granularity description; a stage emotion feature generating unit, configured to perform stage emotion recognition on each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector to obtain a sequence distribution of the stage emotion recognition semantic coding vector; An emotion temporal fluctuation compensation coding unit, used for performing emotion temporal fluctuation context coding based on dynamic compensation on the sequence distribution of the stage emotion recognition semantic coding vector to obtain a stage emotion salient coding vector; The emotion recognition result generating unit is used to obtain the emotion recognition result based on the stage-by-stage emotion significant coding vector.
[0007] According to another aspect of the present application, there is provided a large model-based intelligent question answering method, which includes: Obtain a question description input by a user; input the question description into a large model-based dialogue engine to obtain a reply text; obtain a re-question text description input by the user for the reply text; perform emotion recognition on the re-question text description to obtain an emotion recognition result; in response to the emotion recognition result being unsatisfactory, generate a pop-up window prompt for whether to manually reply; The step of performing emotion recognition on the re-question text description to obtain an emotion recognition result includes: Sentence processing and semantic encoding are performed on the re-question text description to obtain a sequence distribution of semantic encoding vectors of the re-question sentence granularity description; Performing stage emotion recognition on each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector to obtain a sequence distribution of the stage emotion recognition semantic coding vector; Performing emotion temporal fluctuation context encoding based on dynamic compensation on the sequence distribution of the emotion recognition semantic encoding vector of the stage to obtain a stage-specific emotion salient encoding vector; Based on the stage-specific emotion salient coding vector, the emotion recognition result is obtained.
[0008] Compared with the prior art, the large-model-based intelligent question-answering system and method provided by the present application first obtains the question description input by the user, then inputs it into the large-model-based dialogue engine to generate a reply text, then obtains the user's re-question text description for the reply text, and performs emotion recognition on the re-question text description to determine whether the user is satisfied. If the emotion recognition result shows that the user is not satisfied, a prompt pop-up window is generated to ask whether manual intervention is required to reply. In this way, while the intelligent dialogue engine efficiently handles user questions, it can also use emotion recognition technology to perceive user satisfaction, ensuring that when the user is dissatisfied, the option of transferring to manual service is provided, thereby improving user experience and service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 It is a system block diagram of a large model-based intelligent question-answering system according to an embodiment of the present application.
[0010] Figure 2 This is a block diagram of a user re-questioning emotion recognition module in a large-model-based intelligent question-answering system according to an embodiment of the present application.
[0011] Figure 3 It is a block diagram of a text description sentence encoding unit in a large model-based intelligent question-answering system according to an embodiment of the present application.
[0012] Figure 4 It is a block diagram of the emotion timing fluctuation compensation coding unit in the intelligent question-answering system based on a large model according to an embodiment of the present application.
[0013] Figure 5 It is a flowchart of the intelligent question-answering method based on a large model according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0015] In the current era of rapid development of informatization, obtaining information quickly and accurately has become an important need for people. As an efficient means of information interaction, intelligent question-answering systems are widely used in education, medical care, customer service and other fields with their advantages of instant response and accurate answers. They not only improve the efficiency of information acquisition, but also significantly improve user experience.
[0016] Traditional question-answering systems usually rely on preset rules and a limited knowledge base, which limits their ability to understand the diversity and complexity of natural languages. At the same time, due to the limited size of the knowledge base, the answers provided by traditional systems are often incomplete and may even be outdated or biased. Especially in professional fields such as law, medicine, and academic literature, traditional systems are unable to meet users' needs for the latest and authoritative information. In addition, such systems tend to simply match questions with answers, while ignoring the emotional and psychological needs of users. This one-way information interaction method often fails to give timely responses when users are confused or dissatisfied, thus affecting the overall user experience.
[0017] Based on this, this application proposes an intelligent question-answering system based on a large model, which introduces advanced technologies such as large models, dynamic knowledge integration, incremental learning and emotion recognition, aiming to improve the accuracy of question understanding, enhance the relevance and timeliness of reply content, and provide a more personalized and humane service experience through emotion perception. Specifically, Figure 1 : is a system block diagram of an intelligent question-answering system based on a large model according to an embodiment of the present application. Figure 1 As shown, in the intelligent question-answering system 100 based on the big model, it includes: a user question description data collection module 110, which is used to obtain the question description input by the user; a reply text generation module 120, which is used to input the question description into the dialogue engine based on the big model to obtain the reply text; a user re-question description data collection module 130, which is used to obtain the re-question text description input by the user for the reply text; a user re-question emotion recognition module 140, which is used to perform emotion recognition on the re-question text description to obtain an emotion recognition result; a pop-up window triggering module 150, which is used to generate a pop-up window prompt whether to manually reply in response to the emotion recognition result being unsatisfactory.
[0018] In an embodiment of the present application, the user question description data acquisition module 110 is used to obtain the question description entered by the user. It should be understood that the question description entered by the user generally includes the core content of the question, background information and some limiting conditions. Specifically, the core content of the question is the specific problem that the user wants to understand or solve, which may be the understanding of a certain concept, seeking a certain operation method or requesting specific information; the background information is to help better understand the user's intentions so that the system can better locate the core of the problem; the question description may also include the user's special requirements or restrictions for the answer, such as the desire to obtain the latest data, an explanation of a specific field, or a concise answer. In general, by analyzing the question description provided by the user, the system can accurately understand the user's needs and provide targeted responses.
[0019] In an embodiment of the present application, the reply text generation module 120 is used to input the question description into a dialogue engine based on a large model to obtain a reply text. It should be understood that a large model refers to a deep learning model with a huge number of parameters, especially in the field of natural language processing (NLP). Such models are usually trained based on large-scale data sets and can capture complex patterns, semantic relationships, and contextual information in the language. For example, BERT, GPT series, T5, etc. under the Transformer architecture are all well-known large-scale pre-trained language models. The large model can learn a general language representation from a large amount of unlabeled text data through a self-supervised learning method, and then adapt to specific tasks through fine-tuning. In particular, the powerful representation ability and extensive knowledge coverage of the large model can integrate relevant knowledge bases such as relevant laws and regulations, guidelines, literature, textbooks, classic cases, etc., so that it can give reasonable answers on various topics and better adapt to the complexity of natural language. The present application inputs the question description into a dialogue engine based on a large model to use the excellent context perception and language understanding capabilities of the large model to understand the user's question intention, and output a reply text based on this. For example, when a user inputs "how to relieve stress", the model will understand that the user is looking for ways to relieve stress, and then generate a response based on the knowledge learned, such as "You can relieve stress through exercise, listening to music, communicating with friends, etc." The content of the response text depends on the user's question description and the topic it involves. In detail, a clear answer will be given to the core questions in the question description, such as providing specific facts, data or solutions; for questions that require further explanation, the response text will include detailed explanations, step-by-step instructions or principle introductions to help users understand the problem in depth; in addition, depending on the specific circumstances of the question, the response text may also contain some suggestions or operational guidelines, such as recommending best practices and proposing different ways to solve the problem.
[0020] The following is a detailed description of a specific implementation process of "inputting the question description into a dialogue engine based on a large model to obtain a response text": Before inputting the user's question description into the big model, data preprocessing is an essential step. This is because the questions input by users may come from various channels, with diverse formats and contents. For example, some questions may be input by users through voice, some may be text recognized from pictures, and some may be directly in text form. First, the inputs in different formats must be uniformly converted into text format. In addition to format unification, the text also needs to be noise cleaned. The text input by users often contains some noise information such as special characters, extra spaces, and garbled characters. These noise information will interfere with the big model's understanding of the problem and affect the accuracy of the response. For example, the question description text copied and pasted by users from some web pages may contain HTML tags, which have no practical meaning for the big model and need to be removed. At the same time, extra spaces and garbled characters will also affect the semantic expression of the text and must be cleaned. The cleaned text can more clearly express the user's question intention and provide a good foundation for subsequent processing. In addition, for some languages or specific big models, word segmentation is also necessary. In Chinese, there are no obvious signs separating words. Word segmentation can be used to break sentences into meaningful words, which helps the big model understand the semantics more accurately. For example, after “how to learn programming” is segmented into “how”, “learning” and “programming”, the big model can more clearly identify the meaning of each word and the grammatical relationship between them.
[0021] Next, you need to select a suitable large model and adapt it, which is the key to ensuring system performance. There are many large models available on the market, and different models have different characteristics and applicable scenarios. General large models, such as the GPT series, have a wide range of knowledge coverage and strong language understanding capabilities, and are suitable for handling various types of daily problems. Models optimized for specific fields perform better in professional knowledge and specific tasks. If it is a medical question-and-answer system, you need to choose a model that performs well in medical knowledge and language understanding, so that you can answer users' medical questions more accurately. For legal question-and-answer systems, you can use a large number of legal cases and regulatory texts to fine-tune the pre-trained large model. Through fine-tuning, the model can better adapt to the professional terms and semantic expressions in the legal field, and improve the accuracy and professionalism of answering legal questions. At the same time, loading the selected large model into the system environment and deploying it also needs to consider many factors. The system's computing resources, memory, concurrent processing capabilities, etc. will affect the operating efficiency and stability of the model. If the computing resources are insufficient, the model's inference speed may slow down or even freeze; insufficient memory may cause the model to fail to load or run normally. Therefore, during the deployment process, reasonable resource allocation and optimization are required based on actual conditions.
[0022] Then, the preprocessed question description needs to be converted into a numerical representation that the big model can understand, that is, input encoding processing is performed, which is the basis for the big model to process text. First, the word segmenter corresponding to the big model is used to split the question text into tokens, and each token is mapped to a corresponding index. This process is like giving each token a unique identity to facilitate the big model to identify and process. For example, "hello" may be split into two tokens, "you" and "good", and then mapped to specific numerical indexes respectively. In order to help the big model better understand the structure and boundaries of the input, some special markers, such as start markers and end markers, need to be added to the encoded token sequence. These special markers are like "navigation signs" of the text, allowing the big model to clearly know the starting and ending positions of the input. Finally, the encoded token index is input into the embedding layer of the model to convert it into the corresponding word vector representation. The word vector contains the semantic information of the token and is the basis for the subsequent processing of the big model. By converting tokens into vectors, the big model can perform more in-depth analysis and processing of text in high-dimensional space.
[0023] Next, the encoded question needs to be input into the large model for reasoning, which is the core of the whole process. Large models usually use attention mechanisms to capture the relationship between different words in the input text. The attention mechanism allows the model to determine which words are more important to the current processing step. By calculating the attention score, the model can process long texts and complex semantics more effectively. For example, when processing a long article, the attention mechanism can help the model focus on key sentences and words and ignore irrelevant information. The input word vector will pass through multiple neural network layers of the large model in turn for calculation and feature extraction. Each layer will perform nonlinear transformation on the input and gradually abstract higher-level semantic information. After multiple layers of calculation, the model outputs a probability distribution about the words, indicating the possibility of each word as the next reply word. This probability distribution is generated by the model based on the input question and its own learning experience, and it reflects the model's judgment on the possibility of different words appearing in the reply.
[0024] Then, according to the probability distribution output by the large model, it is necessary to select appropriate words to generate reply text, which is a key step in achieving question-answer interaction. There are many methods for generating replies, and greedy search is a simple and efficient method. It selects the word with the highest probability as the next reply word each time until the end tag is generated or the preset maximum length is reached. The advantage of this method is fast speed, but the disadvantage is that it may cause the generated replies to be relatively simple. Because it always selects the most likely word, it lacks a certain diversity. Beam search retains multiple candidate words with high probability at each step to form multiple possible reply paths. Then, at the end of the search, the path with the highest score is selected as the final reply. Beam search can avoid the limitations of greedy search to a certain extent and generate richer and more diverse replies. It is like exploring in multiple possible directions and finally selecting the optimal path. The sampling method randomly samples words from the probability distribution to increase the diversity of replies. By adjusting the sampling temperature parameter, the degree of randomness can be controlled. The higher the temperature, the greater the randomness, and the generated replies may be more diverse, but some unreasonable content may also appear. Therefore, it is necessary to select a suitable generation method according to the specific application scenario and requirements.
[0025] Finally, the generated reply text is post-processed and optimized to further improve the quality and readability of the reply. Grammar correction is very important. Even if the large model has powerful language generation capabilities, some grammatical errors may still occur. Using grammar checking tools or rules to correct the reply can ensure that the reply is grammatically correct. Content screening and filtering are also essential links. It is necessary to filter out sensitive information, wrong information or irrelevant content that may exist in the reply to ensure the security and effectiveness of the reply. If the system is applied to a public question-and-answer platform, the content of the reply must be strictly controlled to avoid bad information. Style adjustment can adjust the style of the reply according to specific application scenarios and user needs. For example, in customer service scenarios, the reply style is usually cordial and friendly, so that users can feel a good service experience; while in academic questions and answers, it must be more rigorous and accurate, in line with academic norms. Through these post-processing and optimization steps, the generated replies can be made more perfect to meet user needs.
[0026] In an embodiment of the present application, the user re-question description data acquisition module 130 is used to obtain the re-question text description input by the user for the reply text. It should be understood that the re-question text description specifically refers to the question or request information further raised by the user after receiving the reply given by the system based on his initial question. Specifically, the user may feel that the initial reply is too broad and needs more specific information, or the user may generate new related questions based on the content of the initial reply. In addition, it may also contain feedback information from the user on the initial reply, for example, "The method you gave is too complicated, is there a simpler one?" Here, "too complicated" indicates that the user is not satisfied with the method provided in the initial reply and expects to get simpler content. By responding to the user's re-questions in a timely manner, especially when it comes to clarifying misunderstandings or providing more detailed information, it can help eliminate the user's doubts and thus improve their overall satisfaction. Moreover, by analyzing the content of the user's re-questions, the system can identify possible deficiencies in the initial answer, such as insufficient information and unclear language expression. Based on these feedbacks, the intelligent question-answering system continues to improve algorithms and knowledge bases to improve service quality and accuracy.
[0027] In the embodiment of the present application, the user re-questioning emotion recognition module 140 is used to perform emotion recognition on the re-questioning text description to obtain an emotion recognition result. It should be understood that the user's re-questioning often contains the true attitude towards the first reply. It may be difficult to judge from the surface of the question whether the user is asking further questions because of insufficient information or expressing dissatisfaction with the reply. Therefore, through emotion recognition, the user's true feelings can be deeply understood, and once dissatisfaction or confusion is detected, the system can take immediate action. However, traditional emotion recognition methods mainly determine emotional tendencies based on preset vocabularies, dictionaries and fixed rule sets. Although this method is more effective for explicitly expressed emotions (such as "happy" and "angry"), it is insufficient when dealing with implicit or indirectly expressed emotions. Due to the lack of in-depth understanding of context and context, traditional systems find it difficult to capture the evolution trend of emotions over time and the logical connection between sentences, which significantly reduces the accuracy and delicacy of emotion recognition.
[0028] Correspondingly, in the user re-question emotion recognition module, the technical concept of the present application is to use an AI-based natural language analysis and processing algorithm to perform sentence processing and semantic encoding on the re-question text description, and then perform stage emotion recognition on the semantic features of the granular description of each re-question sentence, so as to automatically obtain the emotion recognition result by compensating the context representation based on the emotion temporal fluctuations between the semantic encoding features of emotion recognition at each stage. The present application can parse the user's input more carefully, identify implicit or indirectly expressed emotions, and ensure that the emotional features of each sentence can be captured individually and accurately, rather than just relying on surface vocabulary, so that the system can more comprehensively understand the user's true intentions and emotional state, and provide users with a more satisfactory service experience.
[0029] Figure 2 FIG. 1 is a block diagram of a user re-questioning emotion recognition module in a large model-based intelligent question-answering system according to an embodiment of the present application. Specifically, Figure 2 As shown, the user re-question emotion recognition module 140 includes: a text description sentence encoding unit 141, which is used to perform sentence processing and semantic encoding on the re-question text description to obtain a sequence distribution of re-question sentence granularity description semantic encoding vectors; a stage emotion feature generation unit 142, which is used to perform stage emotion recognition on each re-question sentence granularity description semantic encoding vector in the sequence distribution of the re-question sentence granularity description semantic encoding vector to obtain a sequence distribution of the stage emotion recognition semantic encoding vector; an emotion temporal fluctuation compensation encoding unit 143, which is used to perform emotion temporal fluctuation context encoding based on dynamic compensation on the sequence distribution of the stage emotion recognition semantic encoding vector to obtain a stage emotion salient encoding vector; an emotion recognition result generation unit 144, which is used to obtain the emotion recognition result based on the stage emotion salient encoding vector.
[0030] In the embodiment of the present application, the text description sentence encoding unit 141 is used to perform sentence processing and semantic encoding on the re-question text description to obtain a sequence distribution of semantic encoding vectors of the re-question sentence granularity description. Specifically, Figure 3 FIG. 1 is a block diagram of a text description sentence encoding unit in a large model-based intelligent question answering system according to an embodiment of the present application. Figure 3 As shown, the text description sentence encoding unit 141 includes: a re-question text sentence processing unit 1411, which is used to perform sentence processing on the re-question text description to obtain a sequence distribution of re-question sentence granularity descriptions; a re-question text sentence granularity semantic encoding unit 1412, which is used to use a Bi-RNN-based description semantic encoder to semantically encode each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity descriptions to obtain a sequence distribution of the re-question sentence granularity description semantic encoding vectors.
[0031] In an embodiment of the present application, the re-question text sentence processing unit 1411 is used to perform sentence processing on the re-question text description to obtain a sequence distribution of the re-question sentence granularity description. Accordingly, considering the complex sentence structure in natural language, different sentences may contain different emotional information. A user's re-question may consist of multiple sentences, and the content and emotion expressed in each sentence may be different. Therefore, in order to be able to understand and analyze the semantics and emotional information contained in the user's re-question text description more carefully, in the technical solution of the present application, the re-question text description is sentence processed to divide the long text into sentences, and obtain a sequence distribution of the re-question sentence granularity description. In this way, a more accurate emotional analysis can be performed on each individual sentence, so as to better understand the user's specific emotional tendencies in different parts, avoid ignoring subtle emotional differences, and thereby improve the accuracy of emotion recognition.
[0032] In an embodiment of the present application, the re-question text sentence granularity semantic encoding unit 1412 is used to use a description semantic encoder based on Bi-RNN to semantically encode each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity description to obtain a sequence distribution of the re-question sentence granularity description semantic encoding vector. It should be understood that each re-question sentence contains rich semantic information, and these semantic information contain the user's emotional information. Based on this, in the technical solution of the present application, each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity description is semantically encoded to capture the relationship between words and the overall meaning of the sentence, and obtain the sequence distribution of the re-question sentence granularity description semantic encoding vector. In this way, the logical relationship and emotional flow between sentences can be better understood. In particular, in an example of the present application, a description semantic encoder based on Bi-RNN is used to semantically encode each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity description to obtain the sequence distribution of the re-question sentence granularity description semantic encoding vector. Those skilled in the art should know that Bi-RNN consists of a forward RNN and a backward RNN, the forward RNN starts processing from the beginning of the sequence, and the backward RNN starts processing from the end of the sequence. This enables Bi-RNN to simultaneously consider the preceding and following information of each sentence in the question granularity description sequence, which ensures that the encoding of each word or phrase contains the surrounding context information, which is very important for correctly understanding user intent and emotion.
[0033] The following is a detailed description of an implementation process of "using a description semantic encoder based on Bi-RNN to semantically encode each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity description to obtain the sequence distribution of the semantic encoding vector of the re-question sentence granularity description": From the data source, the sequence distribution data of the re-question granularity description is complex and diverse, and pre-processing is required to meet the requirements of subsequent operations. In the framework of natural language processing, word segmentation is the primary and basic task. There is no space between words in Chinese sentences as a natural separator, which brings challenges to word segmentation. At this time, tools such as jieba word segmentation can be applied. Taking the common re-question sentence "The computer I just bought runs very slowly, how can I solve it?" as an example, the jieba word segmentation will split it into independent word units, namely "I", "just", "bought", "of", "computer", "run", "speed", "very slow", "should", "how", "solve". Although English text can usually be simply split by space, abbreviations such as "isn't" and "let's" still need special processing to correctly parse them into "is", "not" and "let", "us", to ensure the accuracy of vocabulary segmentation. After completing word segmentation, the next step is to build a word list. Collect all the words that appear in the re-question granularity description and assign a unique index to each word. Assume that "computer" is assigned index 100 and "run" is assigned index 101. These indexes will become important identifiers for computer recognition and processing of vocabulary in subsequent operations. At the same time, in order for the model to uniformly process re-questions of different lengths, padding and truncation operations are essential. According to the pre-set maximum length standard, if the actual length of the sentence is insufficient, it is padded with specific filler words (such as "PAD"); if it exceeds the maximum length, it is truncated. For example, if the maximum length is set to 40, and a re-question sentence has only 30 tokens after word segmentation, then 10 "PAD" are added to the end of the sentence; if the sentence has 45 tokens, only the first 40 tokens are retained to ensure that the length of the granular description of all re-questions is consistent, which is convenient for subsequent processing.
[0034] After data preprocessing is completed, word embedding processing is required. The core purpose of word embedding is to convert words in the text into low-dimensional real vectors, which can effectively capture the semantic information of words. At present, Word2Vec, GloVe and FastText are common word embedding methods. Taking Word2Vec as an example, it uses a neural network model to build word vector representation by learning the co-occurrence relationship between words in a large amount of text data. In a large amount of text, if "computer" and "software" appear frequently at the same time, according to the Word2Vec principle, their corresponding vectors will be close in spatial position, which intuitively reflects the semantic relevance between the two. GloVe generates word vectors based on global word frequency statistics, and determines the vector representation for each word by analyzing the co-occurrence frequency of words in the entire corpus. FastText not only considers complete words, but also pays attention to the subword information inside the word, which makes it have obvious advantages in processing unregistered words. After selecting the word embedding method, each word in the granular description of the question sentence is mapped to the corresponding word vector. If a 200-dimensional word vector is used, the word "computer" will be converted into a real number vector of length 200, and each dimension of the vector carries information related to the semantics of "computer". After the word embedding operation, the granular description of the question sentence originally composed of words is converted into an ordered sequence of word vectors, which is fully prepared for subsequent model processing.
[0035] Then, the construction of Bi-RNN model is a key step in semantic encoding. Bi-RNN consists of two recurrent neural networks, forward and reverse. This unique structure enables it to simultaneously obtain past and future information in the sequence, thereby capturing semantics more comprehensively and accurately. In practical applications, optimized variants of RNN can be selected, such as LSTM (Long Short-Term Memory Network) and GRU (Gated Recurrent Unit). Taking LSTM as an example, it introduces input gate, forget gate and output gate to finely control the flow and memory of information. The input gate determines the proportion of current input information entering the memory unit, the forget gate controls which information in the memory unit needs to be retained or discarded, and the output gate determines which information in the memory unit is used for the calculation of the current time step. When constructing the Bi-RNN model, it is crucial to reasonably determine key parameters such as the size and number of hidden layers. The size of the hidden layer directly affects the model's ability to learn complex features. A larger hidden layer can learn more complex semantic features, but it will increase the amount of calculation and training time; the number of layers is related to the depth and expression ability of the model. Although increasing the number of layers can improve the ability to learn advanced semantic features, it may also cause overfitting problems. Therefore, it is necessary to scientifically and reasonably select these parameters based on the actual situation such as the amount of data, the complexity of the task, etc. Assume that the hidden layer size is set to 150, which means that the hidden state of each time step is a vector of length 150, which can carry rich semantic information.
[0036] After the model is built, the core process of semantic encoding is entered. The granular description of the re-question sentence that has been preprocessed and word-embedded is input into the Bi-RNN model for semantic encoding. During forward propagation, the forward RNN starts processing from the first word vector of the sequence. For example, for the sentence "The newly downloaded APP always crashes, what's going on?", the word vector of "new" is processed first, and the hidden state is calculated based on the model parameters and the current state. This hidden state not only contains the semantics of "new", but also integrates the model's understanding of the initial input information (in the first time step, the hidden state is often initialized to a zero vector). Then, this hidden state is input into the next time step together with the word vector of the next word "download", and the hidden state is updated. This is continuously passed backward, and the hidden state gradually accumulates the previous information of the sentence. At the same time, back propagation is carried out synchronously. The reverse RNN starts to process backward from the last word vector of the sequence, first processes the word vector of "what's going on", calculates the reverse hidden state, and then processes other word vectors forward in sequence. The reverse hidden state then carries the information of the subsequent words of the sentence. At each time step, the forward and reverse hidden states are merged to obtain the final hidden state. The most common merging method is concatenation, which is to connect the forward and reverse 150-dimensional hidden state vectors in sequence to form a 300-dimensional vector, which fully reflects the semantics of the current word in the sentence.
[0037] For each granular description of the re-question sentence, after being processed by the Bi-RNN model, a corresponding vector can be obtained. This vector can reflect the contextual information of each granular description of the re-question sentence. By integrating the vectors corresponding to the granular description of each re-question sentence, the sequence distribution of the semantic encoding vector of the granular description of the re-question sentence can be obtained. The rich semantic information contained in this sequence distribution can provide key data support for subsequent emotion recognition tasks.
[0038] In the embodiment of the present application, the stage emotion feature generation unit 142 is used to perform stage emotion recognition on each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector to obtain the sequence distribution of the stage emotion recognition semantic coding vector. Specifically, in the embodiment of the present application, the stage emotion feature generation unit is used to: input each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector into the stage emotion recognizer based on the classifier to obtain the sequence distribution of the stage emotion recognition semantic coding vector. Accordingly, considering that each re-question sentence granularity description semantic coding vector contains different emotional components or intensities, and even in the same re-question text, the emotions of each sentence may be different. Therefore, in order to more accurately capture the specific emotions expressed by each sentence, so as to accurately capture these subtle emotional changes, in the technical solution of the present application, each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector is input into the stage emotion recognizer based on the classifier to obtain the sequence distribution of the stage emotion recognition semantic coding vector. It should be understood that the classifier-based stage emotion recognizer is a model that uses a classification algorithm to map input text features or vectors to different emotion categories or stages. It usually takes a certain representation of the text (such as a semantic encoding vector describing the granularity of a re-question sentence) as input, processes it through a trained classifier, and outputs the corresponding emotion category, such as happiness, sadness, anger, surprise, etc., or a more detailed emotional stage. In this way, one or more emotion labels (such as happiness, anger, confusion, etc.) can be assigned to each re-question sentence, forming a sequence distribution of the stage emotion recognition semantic encoding vector, which helps to fully understand the user's immediate emotional state.
[0039] In the embodiment of the present application, the emotion temporal fluctuation compensation coding unit 143 is used to perform emotion temporal fluctuation context coding based on dynamic compensation on the sequence distribution of the stage emotion recognition semantic coding vector to obtain the stage emotion significant coding vector. Specifically, Figure 4 FIG. 1 is a block diagram of an emotion temporal fluctuation compensation coding unit in an intelligent question-answering system based on a large model according to an embodiment of the present application. Figure 4As shown, the emotion temporal fluctuation compensation coding unit 143 includes: a stage emotion semantic hub extraction subunit 1431, which is used to perform sequence hub extraction on the sequence distribution of the stage emotion recognition semantic coding vector to obtain the stage emotion recognition semantic hub feature vector; a stage emotion significant node-hub complementary information coding subunit 1432, which is used to respectively calculate the complementary significant features of each stage emotion recognition semantic coding vector in the sequence distribution of the stage emotion recognition semantic hub feature vector and the stage emotion recognition semantic coding vector to obtain the sequence distribution of the stage emotion recognition significant node-hub complementary information embedded coding vector; an emotion stage feature fusion subunit 1433, which is used to fuse the stage emotion recognition semantic hub feature vector and the sequence distribution of the stage emotion recognition significant node-hub complementary information embedded coding vector to obtain the stage emotion significant coding vector.
[0040] It should be understood that the emotions in the re-question text do not exist in isolation, but show stage-by-stage dynamic changes with the order of sentences. For example, the user may first express slight dissatisfaction, and as the explanation deepens, the emotions gradually become stronger. Moreover, the emotion of each re-question sentence depends not only on the content of the current sentence, but also on the content of the previous and subsequent conversations. Therefore, in order to capture the ups and downs of emotions in the text, the system can more accurately understand the evolution of user emotions to fully grasp this dynamic nature. The present application performs emotional time series fluctuation context encoding based on dynamic compensation on the sequence distribution of the emotional recognition semantic coding vector of the stage to obtain a stage-by-stage emotional salient coding vector. In particular, this method can not only accurately capture the complex relationship between the emotional semantics of each stage, but also dynamically adjust the semantic differences between different features based on the core semantic information, thereby reflecting the development trend of emotions throughout the conversation.
[0041] Specifically, in the embodiment of the present application, the stage emotion semantic hub extraction subunit is used to: multiply each stage emotion recognition semantic coding vector in the sequence distribution of the stage emotion recognition semantic coding vector by the stage emotion recognition semantic weight matrix respectively, and then add them by position with the stage emotion recognition semantic bias vector to obtain the sequence distribution of the stage emotion recognition semantic modulation coding vector; multiply each stage emotion recognition semantic modulation coding vector in the sequence distribution of the stage emotion recognition semantic modulation coding vector by the transformation score vector to obtain the sequence distribution of the stage emotion recognition semantic transformation score value; input the sequence distribution of the stage emotion recognition semantic transformation score value into function to obtain the sequence distribution of the semantic transformation weight values of the stage emotion recognition; weighted fusion of each group of corresponding stage emotion recognition semantic transformation weight values and stage emotion recognition semantic coding vectors in the sequence distribution of the stage emotion recognition semantic transformation weight values and the sequence distribution of the stage emotion recognition semantic coding vector to obtain the semantic hub feature vector of the stage emotion recognition. The above process can be expressed as: ;in, represents the sequence distribution of the emotion recognition semantic encoding vector at the said stage, Express Perform sequence hub extraction, They represent the first, second, and third order of sequence distribution of the semantic encoding vector of emotion recognition in the stage and The semantic encoding vector of emotion recognition in the stage, represents matrix multiplication, express The corresponding stage emotion recognition semantic weight matrix, express The corresponding stage emotion recognition semantic bias vector, represents the transformed score vector, The first The score value of semantic temporal transformation of emotion recognition in each stage, represents the normalized exponential function, The first in the sequence distribution of the weight value of the semantic transformation of emotion recognition The weight value of emotion recognition semantic transformation in each stage, express The number of vectors in Representation stage semantic hub feature vector for emotion recognition.
[0042] It should be understood that the sequence distribution of the stage emotion recognition semantic encoding vector contains a lot of complex and redundant information, which may interfere with the extraction of key emotion features. Through sequence hub extraction, the most representative hub features can be screened out from the complex input sequence, reducing unnecessary information interference and reducing the complexity of subsequent processing. Specifically, when processing a re-question text containing multiple sentences, each sentence generates a corresponding stage emotion recognition semantic encoding vector, which contains rich semantic details, but not all details play a key role in judging the overall emotion. The sequence hub extraction operation can help extract the core content so as to more accurately identify user emotions. Moreover, the generated stage emotion recognition semantic hub feature vector can reflect the sequence dependency of the initial feature distribution, that is, the relationship between features under different sentence orders. This allows the system to not only consider the emotional features of a single sentence when judging user emotions, but also comprehensively consider the coherence and evolution of emotions in the entire re-question text.
[0043] Specifically, in the embodiment of the present application, the stage emotion significant node-hub complementary information encoding subunit is used to: calculate the complementary information of each stage emotion recognition semantic encoding vector in the sequence distribution of the stage emotion recognition semantic encoding vector relative to the stage emotion recognition semantic hub feature vector to obtain the sequence distribution of the stage emotion recognition node-hub complementary information embedded encoding vector, and the process can be expressed as: ;in, The first in the sequence distribution of the semantic encoding vector of emotion recognition in the representation stage The semantic encoding vector of emotion recognition in the stage, express Activation function, represents point convolutional coding, represents the first weight matrix, represents the second weight matrix, yes The corresponding stage emotion recognition semantic transformation encoding vector, The semantic pivot transformation feature vector of emotion recognition in the representation stage, Indicates taking the absolute value, It means to subtract by position point. The first sequence distribution of the node-hub difference vector representing the stage of emotion recognition The node-hub difference vector of emotion recognition in each stage, The sequence distribution of the node-hub complementary information embedding encoding vector in the expression stage The node-hub complementary information of emotion recognition in each stage is embedded into the encoding vector; The complementary information embedding coding vectors of each stage emotion recognition node-hub complementary information in the sequence distribution of the stage emotion recognition node-hub complementary information embedding coding vector are respectively marked with complementary information based on attention to obtain the sequence distribution of the attention weights of the stage emotion recognition node complementary information. The process can be expressed as: ;in, The sequence distribution of the node-hub complementary information embedding encoding vector in the expression stage The node-hub complementary information of emotion recognition in the first stage is embedded in the encoding vector, represents matrix multiplication, represents the complementary information transformation score vector, yes The corresponding stage emotion recognition node-hub complementary information weight matrix, express The corresponding stage emotion recognition node-hub complementary information bias vector, is the first in the sequence distribution of the emotion recognition attention value Attention value of emotion recognition in each stage, represents the exponential function value with the natural constant e as the base, The first in the sequence distribution of the attention weights of the complementary information of the emotion recognition nodes Attention weights of complementary information of emotion recognition nodes at each stage; Based on the sequence distribution of the attention weights of the complementary information of the emotion recognition nodes at the stage, the sequence distribution of the emotion recognition node-hub complementary information embedded coding vector at the stage is modulated by attention to obtain the sequence distribution of the emotion recognition significant node-hub complementary information embedded coding vector at the stage. The process can be expressed as: ;in, and They represent the first, second and third nodes in the sequence distribution of the node-hub complementary information embedding encoding vector of the stage emotion recognition. The node-hub complementary information of emotion recognition in the first stage is embedded in the encoding vector, and They represent the first, second and third order of attention weights of complementary information of emotion recognition nodes in the sequence distribution. Attention weights of complementary information of emotion recognition nodes in each stage, Represents the sequence distribution of the embedding encoding vectors of the significant node-hub complementary information for emotion recognition at the said stage.
[0044] It should be understood that although the stage emotion recognition semantic hub feature vector can capture the overall emotion pattern, in a complex re-question text, each stage emotion recognition semantic encoding vector may contain unique emotion details. The sequence distribution of the stage emotion recognition node-hub complementary information embedding encoding vector obtained by calculating complementary information can supplement these unique emotion details. Moreover, when considering each stage emotion recognition semantic encoding vector separately, the intrinsic connection between them and the overall emotion may not be obvious. By calculating complementary information, these hidden relationships can also be revealed, thereby helping the model to understand the user's emotions more comprehensively and deeply.
[0045] Accordingly, considering that in the sequence distribution of the stage emotion recognition node-hub complementary information embedded coding vector, each vector contains a large amount of complementary information from different stage emotion recognition nodes and stage emotion hubs. These information sources are extensive and may involve multiple aspects of user emotional expression, such as language expression, context association, etc. If these complicated information are not distinguished and processed, it will make the subsequent emotion analysis difficult and may even lead to wrong emotion judgment. The present application can accurately evaluate the importance of node-hub complementary information by performing an attention-based complementary information significant identification operation on the sequence distribution of the stage emotion recognition node-hub complementary information embedded coding vector, so that higher weights can be assigned to important stage emotion recognition nodes, so that the model will pay more attention to this part of information in subsequent analysis, thereby improving the accuracy of the model's understanding of the user's emotions. Here, the allocation process of attention weight is a soft selection process, and lower weights will be assigned to those nodes that are not very relevant to the user's emotion recognition task. This is equivalent to screening and filtering a large amount of emotional information, removing those interfering factors, which can avoid unnecessary interference to the overall analysis results.
[0046] It should be understood that in the process of emotion recognition of the user's re-question text processed by the intelligent question-answering system, the stage emotion recognition node-hub complementary information embedding coding vector contains a lot of feature information about the user's emotional expression. However, not all these features are equally important for accurately judging user emotions. Through the sequence distribution of the attention weights of the stage emotion recognition node complementary information, the stage emotion recognition node-hub complementary information embedding coding vector is modulated to focus on key information. Specifically, when performing attention modulation, the features corresponding to those high-weight vectors will be enhanced, while the features corresponding to those low-weight vectors will be suppressed, which enables the model to focus on processing important information that is closely related to user emotions, while effectively resisting the misleading of emotion recognition by irrelevant factors, which is conducive to improving the accuracy of emotion recognition.
[0047] Finally, the sequence distribution of the emotion recognition semantic hub feature vector and the emotion recognition significant node-hub complementary information embedding coding vector is fused to obtain the emotion significant coding vector. The above process can be expressed as: ;in, The semantic pivot feature vector of emotion recognition in the representation stage, represents the sequence distribution of the embedding encoding vector of the significant node-hub complementary information of emotion recognition in the said stage, Indicates cascade operation, Represents the significant encoding vector of the stage emotion.
[0048] It should be understood that the semantic hub feature vector of stage emotion recognition can reflect the sequence dependency and pattern of the initial feature distribution, as well as the invariance characteristics contained in the feature sequence, which provides key information for grasping the global trend of emotions; the stage emotion recognition salient node-hub complementary information embedding encoding vector can highlight the actual importance of each stage emotion node in the sequence after attention modulation, while retaining the spatial distribution characteristics of the original features, which means that it can capture various local details in the user's emotional expression. By fusing these two feature data, the global and local information of the original feature distribution can be integrated. This integration avoids the isolated processing of information, allowing the global information and local information to complement and reinforce each other, and can provide richer and more accurate feature representations, thereby providing a more reliable basis for subsequent user emotion recognition results.
[0049] In the embodiment of the present application, the emotion recognition result generation unit 144 is used to obtain the emotion recognition result based on the stage emotion significant coding vector. Specifically, in the embodiment of the present application, the emotion recognition result generation unit is used to: input the stage emotion significant coding vector into the emotion determiner based on the classifier to obtain the emotion recognition result. More specifically, in the embodiment of the present application, the emotion recognition result generation unit is used to: use the fully connected layer of the classifier to fully connect the stage emotion significant coding vector to obtain the stage emotion significant fully connected coding classification feature vector; input the stage emotion significant fully connected coding classification feature vector into the Softmax classification function of the classifier to obtain the emotion recognition result. That is, the stage emotion significant coding vector obtained by the dynamic compensation context coding using the sequence distribution of the stage emotion recognition semantic coding vector is classified, so that the stage emotion significant coding vector is analyzed by the classifier and classified into the predefined emotion category, so as to obtain a specific emotion recognition result, that is, satisfaction or dissatisfaction. In this way, the system can adjust the strategy in time, such as providing more detailed explanations, transferring to manual customer service, etc., to improve user experience and improve user satisfaction.
[0050] In particular, since each stage emotion recognition semantic coding vector in the sequence of the stage emotion recognition semantic coding vectors respectively represents the stage emotion recognition semantics based on local coding semantics, when performing emotion temporal fluctuation context encoding based on dynamic compensation, the emotion recognition semantic hub representation under the local semantic space will have a micro dynamic compensation-macro aggregation deviation relative to the global semantic space domain, thereby causing a dynamic deviation in the mapping of features to class target probabilities when the stage emotion salient coding vector is input into the classifier-based emotion determiner, thereby reducing the accuracy of the obtained emotion recognition results.
[0051] Preferably, inputting the stage-specific emotion significant encoding vector into a classifier-based emotion determiner to obtain the emotion recognition result comprises: Determine the label probability value corresponding to each emotion recognition result label obtained by inputting the stage emotion significant encoding vector based on the emotion determiner of the classifier , and calculate the probability value of each label The square root of the sum of the squares of is used to obtain the significant probability value of the stage emotion. The process can be expressed as: ;in, Represents the label probability value corresponding to each emotion recognition result label, Indicates the total number of emotion recognition result labels, Indicates the significant probability value of stage emotions; The characteristic mean of the stage-by-stage emotion significant coding vector is multiplied by the stage-by-stage emotion significant probability value to obtain the stage-by-stage emotion significant statistical field value. This process can be expressed as: ;in, Indicates the significant probability value of stage emotions, represents the feature mean of the stage-specific emotion significant encoding vector, Indicates the significant statistical field value of the stage emotion; Subtract one from the significant statistical field value of the stage emotion and divide it by the significant statistical field value of the stage emotion to obtain the significant partial probability value of the stage emotion. This process can be expressed as: ;in, represents the significant statistical field value of the stage emotion, Indicates the probability value of significant bias of the stage emotion; Calculate the power function of the stage emotion significant coding vector with the stage emotion significant partial probability value as the exponent , and multiply it by the significant partial probability value of the stage emotion to obtain the significant microscopic representation vector of the stage emotion. The process can be expressed as: ;in, It means point multiplication by position. represents the significant encoding vector of the stage emotion, represents the significant partial probability value of the stage emotion, Representation calculation by is the exponential power function, A significant micro-representation vector representing the stage-specific emotion; After multiplying the stage-by-stage emotion significant encoding vector by the stage-by-stage emotion significant partial probability value, an exponential function with a natural constant as the base is calculated to obtain the stage-by-stage emotion significant macroscopic mapping vector. This process can be expressed as: ;in, It means point multiplication by position. represents the significant encoding vector of the stage emotion, represents the significant partial probability value of the stage emotion, represents the exponential function value with the natural constant e as the base, A significant macroscopic mapping vector representing the stage-specific emotion; After calculating the base 2 logarithm of the stage-by-stage emotion-significant micro-representation vector, the logarithm of the stage-by-stage emotion-significant macro-representation vector is weighted and summed with the stage-by-stage emotion-significant macro-mapping vector to obtain an optimized stage-by-stage emotion-significant encoding vector. This process can be expressed as: ;in, and They represent point-by-point multiplication and point-by-point addition, respectively. and represents the weighted hyperparameter, represents the significant micro-representation vector of the stage-specific emotions, represents the logarithmic function value with the natural constant 2 as the base, represents the significant macroscopic mapping vector of the stage emotion, Representing the optimized stage-by-stage emotion saliency encoding vector; The optimized stage-by-stage emotion saliency encoding vector is input into the classifier-based emotion determiner to obtain the emotion recognition result.
[0052] That is, the partial low-order derivatives of the statistical distribution field corresponding to the stage-by-stage emotion salient coding vector are used as non-overlapping macro-feature representation behavior patches of the stage-by-stage emotion salient coding vector, so as to organize the space of different macro-behavior patches under the non-isotropic backbone structure of the stage-by-stage emotion salient coding vector, so as to strengthen the dynamic sensitivity of the long-series micro-complex information distribution of the stage-by-stage emotion salient coding vector to the class probability macro-representation behavior, thereby promoting the iterative dynamic consistency of the class target between the classification target and the extracted features during the feature space-class probability mapping, so as to improve the accuracy of the emotion recognition result obtained by inputting the stage-by-stage emotion salient coding vector into the emotion determiner based on the classifier.
[0053] In summary, the user re-question emotion recognition module 140 is clearly explained, which uses an AI-based natural language analysis and processing algorithm to perform sentence processing and semantic encoding on the re-question text description, and then performs stage emotion recognition on the semantic features of the granular description of each re-question sentence, so as to automatically obtain the emotion recognition result by compensating the context representation based on the emotion temporal fluctuations between the semantic encoding features of emotion recognition at each stage. In this way, the user's input can be analyzed more carefully, and the implicit or indirectly expressed emotions can be identified, ensuring that the emotional features of each sentence can be captured individually and accurately, rather than just relying on surface vocabulary, so that the system can more comprehensively understand the user's true intentions and emotional state, and thus provide users with a more satisfactory service experience.
[0054] In an embodiment of the present application, the pop-up window trigger module 150 is used to generate a pop-up window prompt for whether to manually reply in response to the emotion recognition result being dissatisfied. It should be understood that when the system detects that the user is dissatisfied with the reply content, a pop-up window prompt is used to ask whether manual customer service intervention is needed, which can quickly respond to the user's negative emotions and make the user feel valued and respected. This immediate response helps to alleviate the user's dissatisfaction. Moreover, considering that the intelligent question and answer system is based on a large model, it may not be able to give satisfactory answers to users when faced with some complex, professional or vague questions. The manual customer service has been professionally trained, has richer knowledge and experience, and can conduct in-depth analysis and answers according to the specific situation. Guiding users to choose manual replies through pop-up window prompts can enable users to obtain accurate and comprehensive solutions more quickly, which is conducive to improving the efficiency of problem solving.
[0055] In summary, the big model-based intelligent question-answering system 100 according to the embodiment of the present application is explained, which first obtains the question description input by the user, then inputs it into the big model-based dialogue engine to generate a reply text, then obtains the user's re-question text description for the reply text, and performs emotion recognition on the re-question text description to determine whether the user is satisfied. If the emotion recognition result shows that the user is not satisfied, a prompt pop-up window is generated to ask whether manual intervention is required to reply. In this way, while the intelligent dialogue engine efficiently handles user questions, it can also use emotion recognition technology to perceive user satisfaction, ensuring that when the user is dissatisfied, the option of transferring to manual service is provided, thereby improving user experience and service quality.
[0056] Figure 5 FIG. 1 is a flow chart of an intelligent question answering method based on a large model according to an embodiment of the present application. Figure 5As shown, in the intelligent question-answering method based on the big model, it includes: S110, obtaining a question description input by a user; S120, inputting the question description into a dialogue engine based on the big model to obtain a reply text; S130, obtaining a re-question text description for the reply text input by the user; S140, performing emotion recognition on the re-question text description to obtain an emotion recognition result; S150, in response to the emotion recognition result being unsatisfactory, generating a pop-up window prompt whether to manually reply.
[0057] Here, those skilled in the art can understand that the specific operations of each step in the above-mentioned intelligent question answering method based on a large model have been referred to above. Figures 1 to 4 It has been introduced in detail in the description of the large model-based intelligent question answering system, and therefore, its repeated description will be omitted.
[0058] In summary, the intelligent question-answering method based on the big model of the embodiment of the present application is explained, which first obtains the question description input by the user, then inputs it into the big model-based dialogue engine to generate a reply text, then obtains the user's re-question text description for the reply text, and performs emotion recognition on the re-question text description to determine whether the user is satisfied. If the emotion recognition result shows that the user is not satisfied, a prompt pop-up window is generated to ask whether manual intervention is required to reply. In this way, while the intelligent dialogue engine efficiently handles user questions, it can also use emotion recognition technology to perceive user satisfaction, ensuring that the option of transferring to manual service is provided when the user is dissatisfied, thereby improving user experience and service quality.
Claims
1. An intelligent question-answering system based on a large model, characterized in that: include: A user question description data collection module is used to obtain the question description input by the user; A reply text generation module is used to input the question description into a dialogue engine based on a large model to obtain a reply text; a user re-question description data collection module is used to obtain a re-question text description input by the user for the reply text; a user re-question emotion recognition module is used to perform emotion recognition on the re-question text description to obtain an emotion recognition result; A pop-up window triggering module, used for generating a pop-up window prompt whether to manually reply in response to the emotion recognition result being unsatisfactory; The user then asks the emotion recognition module questions, including: A text description sentence encoding unit, used for performing sentence processing and semantic encoding on the re-question text description to obtain a sequence distribution of semantic encoding vectors of the re-question sentence granularity description; a stage emotion feature generating unit, configured to perform stage emotion recognition on each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector to obtain a sequence distribution of the stage emotion recognition semantic coding vector; An emotion temporal fluctuation compensation coding unit, used for performing emotion temporal fluctuation context coding based on dynamic compensation on the sequence distribution of the stage emotion recognition semantic coding vector to obtain a stage emotion salient coding vector; The emotion recognition result generating unit is used to obtain the emotion recognition result based on the stage-by-stage emotion significant coding vector.
2. The intelligent question-answering system based on a large model according to claim 1, characterized in that: The text description sentence encoding unit includes: A re-question text sentence processing unit, used for performing sentence processing on the re-question text description to obtain a sequence distribution of the re-question sentence granularity description; The re-question text sentence granularity semantic encoding unit is used to use a Bi-RNN-based description semantic encoder to semantically encode each re-question sentence granularity description in the sequence distribution of the re-question sentence granularity description to obtain a sequence distribution of the re-question sentence granularity description semantic encoding vector.
3. The intelligent question-answering system based on a large model according to claim 2, characterized in that: The stage emotion feature generating unit is used to input each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector into a classifier-based stage emotion recognizer to obtain the sequence distribution of the stage emotion recognition semantic coding vector.
4. The intelligent question-answering system based on a large model according to claim 3, characterized in that: The emotion temporal fluctuation compensation coding unit comprises: A stage emotion semantic hub extraction subunit, used for performing sequence hub extraction on the sequence distribution of the stage emotion recognition semantic encoding vector to obtain a stage emotion recognition semantic hub feature vector; The stage emotion salient node-hub complementary information encoding subunit is used to respectively calculate the complementary salient features of the emotion recognition semantic coding vectors of each stage in the sequence distribution of the stage emotion recognition semantic hub feature vector and the stage emotion recognition semantic coding vector to obtain the sequence distribution of the stage emotion recognition salient node-hub complementary information embedded coding vector; The emotion stage feature fusion subunit is used to fuse the sequence distribution of the stage emotion recognition semantic hub feature vector and the stage emotion recognition significant node-hub complementary information embedding coding vector to obtain the stage emotion significant coding vector.
5. The intelligent question-answering system based on a large model according to claim 4, characterized in that: The emotional semantic hub extraction subunit of the stage is used to: Each stage emotion recognition semantic coding vector in the sequence distribution of the stage emotion recognition semantic coding vector is multiplied by the stage emotion recognition semantic weight matrix respectively, and then added by position with the stage emotion recognition semantic bias vector to obtain the sequence distribution of the stage emotion recognition semantic modulation coding vector; Multiplying each stage emotion recognition semantic modulation coding vector in the sequence distribution of the stage emotion recognition semantic modulation coding vector with the transformation score vector to obtain a sequence distribution of the stage emotion recognition semantic transformation score value; Input the sequence distribution of the emotion recognition semantic transformation score values of the stage Function to obtain the sequence distribution of semantic transformation weight values of stage emotion recognition; The sequence distribution of the stage emotion recognition semantic transformation weight values and the sequence distribution of the stage emotion recognition semantic coding vector are weightedly fused with each group of corresponding stage emotion recognition semantic transformation weight values and stage emotion recognition semantic coding vectors to obtain the stage emotion recognition semantic hub feature vector.
6. The intelligent question-answering system based on a large model according to claim 5, characterized in that: The stage emotional salient node-hub complementary information encoding subunit is used to: Calculating the complementary information of the emotion recognition semantic coding vectors of each stage in the sequence distribution of the emotion recognition semantic coding vectors of the stage relative to the emotion recognition semantic hub feature vector of the stage to obtain the sequence distribution of the stage emotion recognition node-hub complementary information embedded coding vectors; The complementary information embedding coding vectors of each stage emotion recognition node-hub complementary information in the sequence distribution of the stage emotion recognition node-hub complementary information embedding coding vectors are respectively marked with complementary information based on attention to obtain the sequence distribution of the attention weights of the complementary information of the stage emotion recognition nodes; Based on the sequence distribution of the attention weights of the complementary information of the emotion recognition nodes at the said stage, the sequence distribution of the emotion recognition node-hub complementary information embedded coding vector at the said stage is modulated by attention to obtain the sequence distribution of the emotion recognition salient node-hub complementary information embedded coding vector at the said stage.
7. The intelligent question-answering system based on a large model according to claim 6, characterized in that: The emotion recognition result generating unit is used to: input the stage-specific emotion significant encoding vector into a classifier-based emotion determiner to obtain the emotion recognition result.
8. The intelligent question-answering system based on a large model according to claim 7, characterized in that: The emotion recognition result generating unit is used to: Using the fully connected layer of the classifier to perform fully connected encoding on the stage-by-stage emotion significant encoding vector to obtain a stage-by-stage emotion significant fully connected encoding classification feature vector; The stage-by-stage emotion significant fully connected encoding classification feature vector is input into the Softmax classification function of the classifier to obtain the emotion recognition result.
9. An intelligent question answering method based on a large model, characterized in that: include: Get the question description entered by the user; Input the question description into a large model-based dialogue engine to obtain a response text; Acquire a re-question text description for the reply text input by the user; perform emotion recognition on the re-question text description to obtain an emotion recognition result; in response to the emotion recognition result being unsatisfactory, generate a pop-up window prompt for whether to manually reply; The step of performing emotion recognition on the re-question text description to obtain an emotion recognition result includes: Sentence processing and semantic encoding are performed on the re-question text description to obtain a sequence distribution of semantic encoding vectors of the re-question sentence granularity description; Performing stage emotion recognition on each re-question sentence granularity description semantic coding vector in the sequence distribution of the re-question sentence granularity description semantic coding vector to obtain a sequence distribution of the stage emotion recognition semantic coding vector; Performing emotion temporal fluctuation context encoding based on dynamic compensation on the sequence distribution of the emotion recognition semantic encoding vector of the stage to obtain a stage-specific emotion salient encoding vector; Based on the stage-specific emotion salient coding vector, the emotion recognition result is obtained.
10. The intelligent question answering method based on a large model according to claim 9, characterized in that: The emotion recognition result is obtained based on the stage-by-stage emotion significant coding vector, including: inputting the stage-by-stage emotion significant coding vector into a classifier-based emotion determiner to obtain the emotion recognition result.
Citation Information
Patent Citations
Method and device for converting intelligent customer service into manual customer service
CN110033281A
Dialogue emotion recognition method and related device
CN114020897A
Emotion recognition method and device, electronic equipment and storage medium
CN115599894A
Business processing method and device based on user emotion recognition, equipment and medium
CN116362777A
Aspect-level sentiment analysis system based on multi-type knowledge fusion
CN118069845A
Cited By
Intelligent vending machine fault early warning system based on cloud platform
CN120048038A