Intelligent dialogue platform based on language model and use method thereof
By adopting a language model-based intelligent dialogue platform in the intelligent dialogue system, using multi-stage intention analysis and context management technology, the shortcomings of existing systems in intention recognition and context management in complex dialogues are solved, and more natural, accurate and emotionally adaptable response is achieved, which significantly improves the user experience.
Patent Information
- Application Number
- CN202510035038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent dialogue system is not accurate enough in complex and multi-round dialogue, imperfect context management, and lacks humanized adaptation in response generation, resulting in poor user experience.
The intelligent dialogue platform based on language model is adopted, including user input processing unit, multi-stage dialogue intention analysis unit, context management and tracking unit, and response generation unit. Through convolutional neural networks, graph neural networks and natural language generation technologies, users' basic intentions, real needs and potential intentions are identified, and context management and response optimization are carried out in combination with user emotional state and environmental data.
It significantly improves the system's intention recognition ability in complex scenarios, the continuity of context management and the humanized adaptability of response content, making the generated response more natural, accurate and emotional, and significantly improves the user experience.
Smart Images

Figure CN119961404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dialogue platforms, and in particular to an intelligent dialogue platform based on a language model and a method for using the platform. Background Art
[0002] With the rapid development of artificial intelligence technology, intelligent dialogue systems based on language models have gradually been applied to various fields, such as online customer service, smart assistants, smart home control, etc. These systems mainly receive user input information, identify user needs and intentions, and generate semantically consistent natural language responses, thereby achieving efficient human-computer interaction. In recent years, breakthroughs in natural language processing (NLP) technology, deep learning models, and multimodal fusion technology have enabled dialogue systems to make significant progress in intent recognition, semantic understanding, and response generation, providing users with a smarter and more natural service experience.
[0003] The intelligent dialogue system of the existing technology has many shortcomings in practical applications. First, the intention recognition is not accurate enough. Traditional methods often rely on a single intention recognition method, which makes it difficult to accurately capture the user's real needs and potential intentions in complex, multi-round dialogues, especially in the case of ambiguous expressions or emotional input. Understanding deviations are prone to occur. Secondly, the context management is imperfect. It is difficult for the system to continuously record and dynamically adjust the dialogue context, resulting in information loss, redundancy or incoherent dialogue. Finally, the response generation lacks humanized adaptation and cannot optimize the tone and content based on the user's emotional state and cognitive load. The generated responses lack situational awareness and are difficult to meet the user's personalized needs, resulting in poor user experience.
[0004] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide an intelligent dialogue platform based on a language model and a method of using the same, which effectively improves the system's intent recognition capability in complex scenarios, the continuity of context management and the humanized adaptability of response content, making the generated responses more natural, accurate and emotional, significantly improving the user experience, and enhancing the intelligence level and adaptability of the dialogue system. Summary of the invention
[0005] The present invention provides an intelligent dialogue platform based on a language model and a method for using the platform.
[0006] The intelligent dialogue platform based on the language model includes a user input processing unit, a multi-stage dialogue intention analysis unit, a context management and tracking unit, and a response generation unit, wherein;
[0007] The user input processing unit receives the user's input information and pre-processes the input information, including speech recognition and text segmentation;
[0008] The multi-stage conversation intention analysis unit analyzes the semantic information of the user input information, identifies the user's basic intention, and identifies the user's needs and potential intentions in combination with the user's historical conversations, the user's emotional state, and environmental perception data, specifically including:
[0009] Basic intent recognition: Analyze user input information through the convolutional neural network (CNN) model to identify the user's basic intent, including query, request, and command;
[0010] Contextual reasoning: Based on the identified basic intent and the user's historical conversations, the real needs of the user can be identified;
[0011] Potential intention recognition: Combine the user's emotional state and environmental perception data to infer the user's hidden potential intention;
[0012] The context management and tracking unit records and manages the conversation context information in real time, and adjusts the context information according to the identified user needs and potential intentions;
[0013] The response generation unit generates a corresponding natural language response based on the recognition results of user needs and potential intentions, combined with the adjusted context information, using natural language generation (NLG) technology based on a language model, and optimizes the response content according to the user's emotional state and cognitive load (the cognitive resource consumption required to process information or complete tasks, that is, the information processing burden borne by the brain in the process of understanding, memory and decision-making).
[0014] Optionally, the user input processing unit includes:
[0015] Information reception: receiving user input information, including voice input information and text input information;
[0016] Speech recognition: When receiving speech input information, a deep neural network (DNN) model is used to convert the speech signal into a text vector;
[0017] Text segmentation: When receiving text input information, the BERT model-based word segmentation technology is used to generate a text vector for each word or word fragment.
[0018] Optionally, the convolutional neural network (CNN) model includes:
[0019] Input feature preparation: The text vectors generated by speech recognition and text segmentation are used as input features;
[0020] Convolution operation: Use multiple convolution kernels of different sizes to convolve the input features to capture semantic information of different sizes;
[0021] Non-linear transformation: The output after the convolution operation is non-linearly transformed through the ReLU activation function;
[0022] Pooling operation: Use the maximum pooling method to reduce the dimension of the convolution features after nonlinear transformation;
[0023] Intent classification: The feature map after the pooling operation is flattened and input into the fully connected layer to generate a prediction score for each intent category. The prediction score is converted into a probability distribution by applying the Softmax activation function.
[0024] Intent recognition and output: According to the probability distribution of Softmax output, the category with the highest probability is selected as the final intent recognition result.
[0025] Optionally, the context reasoning includes:
[0026] Intention reasoning and context association: By combining the user’s basic intention and historical conversation content, a graph neural network (GNN) model is used to obtain the context vector C context ;
[0027] Reasoning about real needs: Based on the obtained context vector C context Reasoning about the real needs of users real .
[0028] Optionally, the potential intention identification includes:
[0029] Emotional state recognition: Identify the emotional state vector E in the generated text vector emotion , including happiness, sadness, and anger;
[0030] Environmental perception data processing: obtain the user's current environmental data (device status, geographic location, time, network status, etc.) and represent it as an environmental feature vector E env ;
[0031] Potential intention inference: Combined with the user's emotional state vector E emotion and the environmental feature vector E env , the user's potential intention is inferred through the multimodal fusion model, and the weighted attention mechanism is used to generate the final potential intention vector R latent ;
[0032] Output potential intent: potential intent vector R latent After classification by the Softmax classifier, the user’s hidden potential intention y is obtained latent .
[0033] Optionally, the context management and tracking unit includes:
[0034] Context information recording: Real-time recording of the user's current conversation input information and the conversation platform's response results, building the current conversation context information C current ;
[0035] Context information storage: Maintain a fixed-length context window W to store the most recent L rounds of conversation information;
[0036] Update the association between user needs and intentions: Dynamically update and adjust the current context information based on the identified user needs and potential intentions. If the user intention changes, the current input information U t Restart the conversation and clean up the old irrelevant context information. If the intent remains continuous, update the most recent state in the context window, including the basic intent and potential intent, to maintain the continuity of the conversation.
[0037] Context state optimization: Optimize the stored context information and extract the core content;
[0038] Final context output: Based on the dynamically updated and optimized context information, the final context C is output final .
[0039] Optionally, the response generation unit includes:
[0040] Basic response generation: Generates preliminary natural language responses based on the user's real needs, potential intentions, and final context;
[0041] Emotion and cognitive load optimization: The tone and expression of the generated preliminary natural language response are optimized by combining the user's emotional state vector and cognitive load to generate the final natural language response.
[0042] Optionally, the basic response generation includes:
[0043] Input feature fusion: The user’s real needs R real , potential intention vector y latent and the final context C final Perform feature fusion to generate the response input vector X input ;
[0044] Language model processing: The fused input vector X input Input into the natural language generation (NLG) model to generate a preliminary natural language response R base .
[0045] Optionally, the emotion and cognitive load optimization includes:
[0046] Emotional state adaptation: based on the user's emotional state vector E emotion , adjust the initial natural language response Rbase Tone of voice and emotional expression;
[0047] Cognitive load adaptation: Combined with the user's cognitive load vector C cogload , optimize the complexity and information content of the initial natural language response after the emotional state adaptation;
[0048] Generate the final response: The response R after sentiment optimization and cognitive load optimization final As the final natural language response.
[0049] The method for using the language model-based intelligent dialogue platform is implemented by the above-mentioned language model-based intelligent dialogue platform, and includes the following steps:
[0050] S1, user input reception and preprocessing: user input information is received by the user input processing unit, and the input information is preprocessed, including speech recognition and text segmentation, to convert the input information into a text vector;
[0051] S2, basic intention and potential intention analysis: The preprocessed user input information is input into the multi-stage dialogue intention analysis unit to analyze the semantic content of the input information, identify the user's basic intention, and combine historical dialogues, the user's emotional state and environmental perception data to infer the user's needs and hidden potential intentions;
[0052] S3, context management and adjustment: records the current conversation content in real time and dynamically updates the context information based on the identified user needs and potential intentions;
[0053] S4, natural language response generation: Based on the identified user needs, potential intents, and adjusted context information, a preliminary natural language response is generated through natural language generation technology based on a language model;
[0054] S5, Emotion and cognitive load optimization: Combine the user's emotional state vector and cognitive load to optimize the tone and content of the generated preliminary natural language response, and finally generate a natural language response that meets the user's emotional state, comprehension ability and needs;
[0055] S6, output response: output the final optimized natural language response to the user, completing the intelligent interaction between the user and the platform.
[0056] Beneficial effects of the present invention:
[0057] The present invention realizes efficient processing and in-depth understanding of user input information through the collaborative work of the user input processing unit, the multi-stage dialogue intention analysis unit and the context management and tracking unit. The user input processing unit can efficiently process voice and text input to ensure the accuracy of input information. The multi-stage dialogue intention analysis unit identifies basic intentions, real needs and potential intentions through convolutional neural network and graph neural network models, and performs reasoning based on the user's emotional state and environmental data, effectively solving the problem of inaccurate intention recognition in traditional dialogue systems in complex situations. In addition, the context management and tracking unit can record, manage and optimize context information in real time, ensuring that the system maintains information consistency and accuracy in multiple rounds of dialogues, dynamically adjusts the context state, reduces information redundancy, and enhances the system's continuous understanding and adaptability to user intentions.
[0058] The present invention, through the response generation unit, utilizes natural language generation technology, combines user needs, potential intentions and contextual information, to generate a natural language response with clear logic, accurate content and semantic coherence. On this basis, the response content is dynamically adjusted in combination with the user's emotional state and cognitive load, which can adapt to the user's emotions and provide a warmer, more interactive or soothing response to meet the needs of users in different emotional states. Cognitive load optimization adjusts the complexity of the response according to the user's comprehension ability, which can simplify information to reduce high cognitive load and provide supplementary explanations to meet the needs of users with low cognitive load. This dual optimization strategy effectively improves the adaptability and humanization of the output, making the response more natural, smooth and easy to understand. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 A schematic diagram of a platform functional unit according to an embodiment of the present invention;
[0061] Figure 2 The figure is a flowchart of the method of using an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it is explained here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0063] It should be noted that the references to "one embodiment", "an embodiment", "an exemplary embodiment", "some embodiments" and the like in the specification indicate that the embodiments described may include specific features, structures or characteristics, but not every embodiment may include the specific features, structures or characteristics. In addition, when a specific feature, structure or characteristic is described in conjunction with an embodiment, it should be within the knowledge of a person skilled in the art to implement such feature, structure or characteristic in conjunction with other embodiments (whether or not explicitly described).
[0064] In general, a term can be understood, at least in part, from its use in context. For example, depending, at least in part, on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending, at least in part, on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0065] like Figure 1 As shown, the language model-based intelligent dialogue platform includes a user input processing unit, a multi-stage dialogue intention analysis unit, a context management and tracking unit, and a response generation unit, wherein;
[0066] The user input processing unit receives the user's input information and pre-processes the input information, including speech recognition and text segmentation;
[0067] The multi-stage conversation intention analysis unit analyzes the semantic information of the user's input information, identifies the user's basic intention, and combines the user's historical conversation, the user's emotional state, and environmental perception data to identify the user's needs and potential intentions, including:
[0068] Basic intent recognition: Analyze user input information through the convolutional neural network (CNN) model to identify the user's basic intent, including query, request, and command;
[0069] Contextual reasoning: Based on the identified basic intent and the user's historical conversations, the real needs of the user can be identified;
[0070] Potential intention recognition: Combine the user's emotional state and environmental perception data to infer the user's hidden potential intention;
[0071] The context management and tracking unit records and manages the conversation context information in real time, and adjusts the context information according to the identified user needs and potential intentions;
[0072] The response generation unit generates corresponding natural language responses based on the recognition results of user needs and potential intentions, combined with the adjusted context information, using natural language generation (NLG) technology based on language models, and optimizes the response content according to the user's emotional state and cognitive load (the cognitive resource consumption required to process information or complete tasks, that is, the information processing burden borne by the brain in the process of understanding, memory and decision-making);
[0073] Through the above content, the intelligence and adaptability of the dialogue system are effectively improved, ensuring that the system can accurately identify the user's basic intentions, real needs and potential intentions, while dynamically adjusting the dialogue context and optimizing the generated response content, making the dialogue more natural, smooth and emotionally adaptive, and able to cope with complex multi-round dialogue scenarios, improve user experience, and reduce information omissions or misunderstandings.
[0074] The user input processing unit includes:
[0075] Information reception: receiving user input information, including voice input information and text input information;
[0076] Speech recognition: When receiving speech input information, a deep neural network (DNN) model is used to convert the speech signal into a text vector, including:
[0077] (1) Feature extraction: Extract the Mel-frequency cepstral coefficient (MFCC) features of the input speech signal x(t), including:
[0078] Framing: Divide the input speech signal x(t) into multiple frames x1, x2, …, x m , the length of each frame signal is 20ms to 40ms;
[0079] Window function: Apply a window function to each frame, expressed as:
[0080] x window [n] = x[n]·w[n]
[0081] Among them, x window [n] is the nth sample value of the speech signal after the window function is applied, w[n] is the window function, and x[n] is the original sample value of the signal;
[0082] Fast Fourier Transform: Perform FFT transformation on each frame of signal to obtain frequency domain representation, which is expressed as:
[0083]
[0084] Among them, X k is the complex number representation in the frequency domain, N is the FFT length, and k is the frequency index;
[0085] Mel frequency transform: The spectrum is transformed through the Mel frequency filter bank to obtain the energy characteristics of the Mel frequency band, which is expressed as:
[0086] M f =∑ k |X k | 2 ·H f (k);
[0087] Among them, H f (k) is the filter function of the Mel filter bank, M f is the energy feature vector after Mel frequency transformation, which is the result of passing the spectrum through the Mel filter bank;
[0088] Logarithmic operation and DCT transformation: The Mel frequency energy is processed logarithmically and the discrete cosine transform is applied, which is expressed as:
[0089]
[0090] Among them, MFCC n is the nth MFCC coefficient, M f [m] is the Mel frequency energy, MFCC n is the nth MFCC coefficient, M is the order of DCT;
[0091] (2) Application of speech recognition model: The extracted MFCC features are modeled using a deep neural network (DNN) model to perform speech-to-text conversion. The DNN processes the input feature X and calculates the i The corresponding vocabulary probability distribution P(w i |x(t i )) Select the most likely word sequence by maximizing the probability output It is expressed as:
[0092]
[0093] Among them, P(w i |x(t i )) is each time step t calculated by the DNN network i The lexical probability of
[0094] Text segmentation: When receiving text input information, the BERT model-based word segmentation technology is used to generate text vectors for each word or word fragment, including:
[0095] (1) Text input processing: When receiving text input information, convert the text T = {t1, t2, ..., t m}Remove extra spaces, punctuation marks, and unnecessary characters;
[0096] (2) Text segmentation: Use the BERT model to segment text, including:
[0097] Text segmentation: The input text T is divided into multiple subwords, and the subword representation of each word is obtained through the WordPiece algorithm;
[0098] BERT input representation: Each subword is converted into an embedding vector through the Embedding layer and input into the BERT model. i Corresponding to an embedding vector e i , expressed as:
[0099] e i =BERTEmbedding(t i );
[0100] Among them, t i is the i-th Token;
[0101] Bidirectional Transformer encoding: The BERT model encodes the input Token through a bidirectional Transformer network. Each Token is processed by multiple Transformer layers to obtain its final context-sensitive representation vector E = [e1, e2, ..., e k ];
[0102] Context representation: The context representation vector of each Token is expressed as:
[0103]
[0104] Among them, e i-1 ,e i ,e i+1 They are the vector representations of the current Token and its context, The context representation vector of each Token calculated by BERT’s bidirectional Transformer encoder;
[0105] Through the above content, different forms of user input, including voice and text, can be processed efficiently and accurately. For voice input, deep neural network (DNN) and Mel-frequency cepstral coefficient (MFCC) technology can accurately convert voice signals into text, greatly improving the accuracy and robustness of speech recognition. For text input, context-sensitive word segmentation and encoding through the BERT model can better understand the meaning of each word in different contexts, thereby improving the semantic parsing ability of the text.
[0106] The Convolutional Neural Network (CNN) model includes:
[0107] Input feature preparation: The text vector generated by speech recognition and text segmentation is used as the input feature, expressed as:
[0108] X=[x1,x2,…,x n ];
[0109] Among them, x1,x2,…,x n represents the vector of the 1st, 2nd, ..., nth input, where n is the total length of the input and X is the input feature;
[0110] Convolution operation: Use multiple convolution kernels of different sizes to convolve the input features to capture semantic information of different sizes, expressed as:
[0111] z k =conv(X,W k )+b k ;
[0112] Among them, conv(X,W k ) represents the convolution operation, z k is the feature map after convolution operation, W k is the weight of the convolution kernel, b k is the bias term, k represents the size of different convolution kernels (k=3 and k=5);
[0113] Non-linear transformation: The output after the convolution operation is transformed non-linearly through the ReLU activation function to increase the expressive power of the model, expressed as:
[0114] a k =max(0,z k );
[0115] Among them, a k is the activated feature map, z k is the convolution result;
[0116] Pooling operation: In order to reduce the dimension of features and prevent overfitting, the maximum pooling method is used to reduce the dimension of the convolution features after nonlinear transformation, which is expressed as:
[0117] p k =maxpool(a k ,2×2);
[0118] Among them, p k is the output after pooling, a k It is the feature map after ReLU activation;
[0119] Intent classification: The feature map after the pooling operation is flattened and input into the fully connected layer to generate the prediction score for each intent category. The prediction score is converted into a probability distribution by applying the Softmax activation function, which is expressed as:
[0120] y=W FC p+b FC ;
[0121] Among them, W FC is the weight matrix of the fully connected layer, b FC is the bias term, y is the predicted intent category score vector;
[0122]
[0123] Among them, P(y i |X) represents the probability that the input X belongs to the i-th category, y i is the output score of the i-th category, y j is the output score of the jth category;
[0124] Intent recognition and output: According to the probability distribution of Softmax output, the category with the highest probability is selected as the final intent recognition result, which is expressed as:
[0125]
[0126] Among them, y pred The identified user basic intent category;
[0127] Through the above content, different contextual information in user input is effectively captured, so that diverse basic intentions (such as queries, requests, commands, etc.) can be identified more accurately, and the ability to understand text input is enhanced. Especially when dealing with complex user intentions, it shows strong adaptability. In addition, the use of ReLU activation function and pooling operation not only reduces the computational complexity of the model and prevents overfitting, but also improves the generalization ability of the model. Through the fully connected layer and Softmax classification, the basic intention of the user can be output with high precision, which optimizes the response efficiency and user experience of the entire dialogue platform.
[0128] Contextual reasoning includes:
[0129] Intention reasoning and context association: By combining the user’s basic intention and historical conversation content, a graph neural network (GNN) model is used to obtain the context vector C context , including:
[0130] (1) Graph embedding: embedding the user’s basic intention y pred Dialogue with History historyConverted into vector representation through word embedding method, each word or phrase corresponds to a node embedding vector It is expressed as:
[0131]
[0132] in, is the feature vector of node i at the initial time, Embedding(w i ) is to transform node w i Transformed into a vector representation embedding function, w i ∈{y pred ,D history} is an element in the node set, including the user's basic intention and historical conversation;
[0133] (2) Information propagation: In a graph neural network, each node is updated based on the information of its neighboring nodes. The representation of each node is propagated (i.e., aggregated) through the information of its neighboring nodes. is the node representation after the t-th propagation, and the update of node i is expressed as:
[0134]
[0135] in, is the feature vector of node i after the tth propagation, AGGREGATE(·) is the aggregation function used to aggregate the information of all neighboring nodes of node i, including summation, averaging, and maximum pooling. is the set of feature vectors of neighboring nodes j connected to node i in the t-1th iteration, (i, j)∈ε means there is an edge between nodes i and j in the graph, W (t) is the weight matrix in the t-th propagation;
[0136] (3) Context vector update: After several rounds of information propagation, the representation of each node They all contain the information of their neighbor nodes and context information. The representation of all nodes is aggregated through the fully connected layer to obtain the overall context vector C context , expressed as:
[0137]
[0138] in, is the final feature vector of node i after T propagation iterations, Pooling is the global pooling function, and v is the set of all nodes in the graph;
[0139] Reasoning about real needs: Based on the obtained context vector C context Reasoning about the real needs of users real , expressed as:
[0140]
[0141] Among them, f(D history ,C context ) is the computational history dialogue D history With the context vector C context The similarity function between (cosine similarity), α i is the weight calculated based on the attention mechanism, and m is the total number of historical conversations;
[0142]
[0143] Through the above content, the information most relevant to current needs can be accurately identified, the key content in historical conversations can be effectively mined, key information can be highlighted, information redundancy can be avoided, and the system can accurately understand the user's real needs. In addition, by calculating the correlation between historical conversations and contexts through similarity functions, continuous tracking and in-depth understanding of user intentions can be maintained in complex, multi-round conversation scenarios, thereby greatly improving the accuracy of reasoning and the responsiveness of the system.
[0144] Potential intent identification includes:
[0145] Emotional state recognition: Identify the emotional state vector E in the generated text vector emotion , including happiness, sadness, and anger, expressed as:
[0146] E emotion =Softmax(W e ·X+b e );
[0147] Where X is the text vector generated by the user input processing unit, W e is the weight matrix of the classifier, which is used to map the text vector to the sentiment category space, b e is a bias term, and the Softma function is used to normalize the result of the linear transformation into a probability distribution, which represents the probability that the text belongs to each sentiment category. emotion is the output emotional state vector;
[0148] Environmental perception data processing: obtain the user's current environmental data (device status, geographic location, time, network status, etc.) and represent it as an environmental feature vector E env , expressed as:
[0149] E env =[f1,f2,...,f m ];
[0150] Among them, f1,f2,...,f mIndicates the 1st, 2nd, ..., mth environmental features (location coordinates, device power, network delay, etc.);
[0151] Potential intention inference: Combined with the user's emotional state vector E emotion and the environmental feature vector E env , the user's potential intention is inferred through the multimodal fusion model, and the weighted attention mechanism is used to generate the final potential intention vector R latent , expressed as:
[0152]
[0153] Among them, W c is the weight matrix of multimodal fusion, b c is the bias term, ‖ represents the vector concatenation operation, q is the number of weighted summation terms, β i is the attention weight, which is calculated by the similarity between the emotional state and the environmental data and is defined as:
[0154] Output potential intent: potential intent vector R latent After classification by the Softmax classifier, the user’s hidden potential intention y is obtained latent , expressed as:
[0155] y latent = argmax p Softmax(W latent ·R latent +b latent );
[0156] Among them, W latent and b latent are the weight and bias of the classifier, y latent is the inferred user potential intention category, p is the index of the emotion or intention category;
[0157] Through the above content, not only does it focus on the user's basic intentions, but it also provides a more comprehensive contextual understanding through emotional states and environmental information, enabling the system to have stronger situational awareness. In addition, the attention mechanism is used to dynamically assign weights to emotional states and environmental features, highlight key information, and reduce redundant calculations, thereby improving the accuracy of reasoning and response speed. It shows significant advantages when dealing with ambiguous or unclear intentions, further enhancing the user experience and service quality of the intelligent dialogue system.
[0158] The context management and tracking unit includes:
[0159] Context information recording: Real-time recording of the user's current conversation input information and the conversation platform's response results, building the current conversation context information Ccurrent , expressed as:
[0160] C current ={(U1,S1),(U2,S2),...,(U t ,S t )};
[0161] Among them, C current is the context information recorded in the current round of dialogue, U1,U2,...,U t is the user's input information in the 1st, 2nd, ..., tth round of dialogue, S1, S2, ..., S t is the response content in rounds 1, 2, ..., t, where t is the number of rounds in the current dialogue;
[0162] Context information storage: Maintain a fixed-length context window W to store the conversation information of the most recent L rounds, ensuring that context information is not lost and the computational overhead is controllable, expressed as:
[0163] C window ={(U t-L+1 ,S t-L+1 ),...,(U t ,S t )};
[0164] Where L is the length of the window;
[0165] Update the association between user needs and intentions: Dynamically update and adjust the current context information based on the identified user needs and potential intentions. If the user intention changes, the current input information U t Restart the conversation and clean up the old irrelevant context information. If the intent remains continuous, update the most recent state in the context window, including the basic intent and potential intent, to maintain the continuity of the conversation.
[0166] Context state optimization: Optimize the stored context information and extract the core content to reduce redundant data and improve system processing efficiency, expressed as:
[0167] C optimized =Extract_Core(C current );
[0168] Among them, C optimized It is the optimized core context information, and Extract_Core represents the process of extracting the core content;
[0169] Final context output: Based on the dynamically updated and optimized context information, the final context C is output final ;
[0170] The core content is the content that can reflect the user's current needs, potential intentions, and key information of the conversation after screening and extraction, including:
[0171] Key information extraction: includes keywords, phrases or sentences in user input that are closely related to basic intent and potential intent, such as query objects, request actions, and command instructions in user input;
[0172] Context-related content: Combine historical conversation records to filter out information that is highly relevant to the current conversation. For example, in the "context window", remove repeated or irrelevant sentences and retain the context that helps understand user needs.
[0173] Topic or purpose-oriented: highlight the user's current conversation goal rather than all historical content. For example, when a user asks about "weather", only information related to location, time, and weather is retained, and irrelevant conversations are removed.
[0174] Core content examples:
[0175] Assume the following user dialogue:
[0176] User A: "I'm going to Shanghai on a business trip tomorrow."
[0177] Platform S: "OK, do you need to book a flight?"
[0178] User A: "Yes, help me check the flight status for tomorrow morning."
[0179] Core content extraction:
[0180] Keywords: "Shanghai", "business trip", "tomorrow morning", "flight";
[0181] Key information: The user's demand is to book a flight ticket for tomorrow morning and the destination is Shanghai;
[0182] Core content means: C optimized ={User demand: book air tickets, location: Shanghai, time: tomorrow morning, intention: check flights};
[0183] Through the above content, the system ensures that the consistency and accuracy of information are maintained in multiple rounds of conversations. It can timely update the context status according to the identified user needs and potential intentions, avoid information redundancy, extract key content, and provide the system with concise and accurate context data. In addition, through core content extraction and context optimization, the unit can effectively reduce computational complexity, improve system response speed and efficiency, make the generated responses more in line with the user's current needs and intentions, enhance the continuity and intelligence of the intelligent dialogue platform, and thus provide users with a more natural and smooth interactive experience.
[0184] The response generation unit includes:
[0185] Basic response generation: Generates preliminary natural language responses based on the user's real needs, potential intentions, and final context;
[0186] Emotion and cognitive load optimization: Combine the user's emotional state vector and cognitive load to optimize the tone and expression of the generated preliminary natural language response to generate the final natural language response;
[0187] Through the above content, we have achieved an accurate understanding of user needs, potential intentions and context, and generated more natural response content that fits the user's emotions and cognitive state. The basic response generation ensures that the system can quickly generate preliminary responses with clear logic and accurate content, while the emotion and cognitive load optimization module further adjusts the tone and complexity of the response so that it can adapt to the user's emotional state and be simplified or expanded according to the user's cognitive load.
[0188] Basic response generation includes:
[0189] Input feature fusion: The user’s real needs R real , potential intention vector y latent and the final context C final Perform feature fusion to generate the response input vector X input , expressed as:
[0190] X input =[R real ‖y latent ‖C final ];
[0191] Among them, ‖ represents the vector concatenation operation;
[0192] Language model processing: The fused input vector X input Input into the natural language generation (NLG) model to generate a preliminary natural language response R base , expressed as:
[0193] R base =NLG_Model(X input );
[0194] Among them, NLG_Model represents the natural language generation model;
[0195] The natural language generation model NLG_Model adopts a sequence-to-sequence model (Seq2Seq), which includes:
[0196] (1) Input feature fusion: The response formed by the input generates the input vector X input ;
[0197] (2) Encoder processes the input vector: The fused input vector X input Encoded as a context vector h, which is used by the decoder to generate a response, expressed as:
[0198] Input: X input ={x1,x2,...,x n}, represents the input sequence feature vector;
[0199] Coding process: h t =Encoder(x t ,h t-1 );
[0200] Among them, h t is the hidden state at step t, x t is the feature vector of the input sequence at time step t;
[0201] Output: context vector H = {h1,h2,...,h T}, used for decoder input;
[0202] (3) Decoder generates preliminary natural language response: The decoder gradually generates a natural language response based on the context vector H output by the encoder, which can be expressed as:
[0203] Input: context vector H output by the encoder, initial input of the decoder (start symbol 〈SOS〉);
[0204] Decoding process (generate response sequence step by step):
[0205] s t =Decoder(y t-1 ,s t-1 ,H);
[0206] y t =Softmax(W·s t + b);
[0207] Among them, s t is the hidden state of the decoder at time step t, y t is the probability distribution of the output word generated in the tth step (normalized by Softmax), W and b are the weight matrix and bias term of the output layer respectively, y t-1 is the word or feature vector output by the decoder at time step t-1;
[0208] Generate stop condition: the decoder outputs the end symbol 〈EOS〉 or reaches the maximum sequence length;
[0209] Output: Preliminary generated response text sequence Rbase ={y1,y2,...,y T};
[0210] (4) Output preliminary natural language response: The text sequence R generated by the decoder base Conduct integration as the final preliminary response output;
[0211] Through the above content, user needs, potential intentions and final contextual information are effectively integrated to ensure that the generated natural language response logic is clear, the content is accurate and the semantics are coherent. The encoder can deeply understand the input features and capture the user's real needs and contextual information, while the decoder gradually generates natural language text that meets the user's intentions. It is not only suitable for multi-round dialogue scenarios, but also can handle complex contextual dependencies, ensuring that the generated response fits the user input and provides high-quality basic content, providing a solid foundation for subsequent emotion and cognitive load optimization, and significantly improving the system's response accuracy and intelligence level.
[0212] Emotional and cognitive load optimization includes:
[0213] Emotional state adaptation: based on the user's emotional state vector E emotion , adjust the initial natural language response R base Tone and emotional expression, including:
[0214] (1) Input: preliminary natural language response R base and the emotional state vector E emotion ;
[0215] (2) Emotion optimization: A weighted adjustment mechanism is used to optimize the generated response according to the emotional state, expressed as:
[0216] R emotion =R base +W e ·E emotion ;
[0217] Among them, R base is the initial generated natural language response, W e is the emotion regulation weight matrix, E emotion is the user's emotional state vector, R emotion The text representation of the response adjusted in combination with the affective state;
[0218] (3) Emotion adjustment strategy: For negative emotions (sadness, anger), adjust the tone to be soothing and gentle, and reduce irritating words; for positive emotions (happiness), adjust the tone to be positive and interactive, and increase encouraging words;
[0219] Cognitive load adaptation: Combined with the user's cognitive load vector C cogload, optimize the complexity and information content of the initial natural language response after the emotional state adaptation, including:
[0220] (1) Input: Response R after emotional state adaptation emotion and the user cognitive load state vector C cogload ;
[0221] (2) Cognitive load optimization: Dynamically adjust the response complexity according to the cognitive load vector and optimize by simplifying or expanding the content, expressed as:
[0222] R final =R emotion ·(1-γ)+Simplify(R emotion )·γ;
[0223] Among them, R final For the final optimized natural language response, R emotion Simplify(R) is the response text after sentiment optimization. emotion ) is a simplified version of the response text (removing redundant information and retaining the core content), γ = Sigmoid (W c ·C cogload +b c ), represents cognitive load weight, with a value range of [0,1], W c and b c are the corresponding weights and biases respectively;
[0224] (3) Cognitive load optimization strategy: For high cognitive load, increase γ, simplify the response content, highlight the core information, and reduce the user's cognitive burden. For low cognitive load, reduce γ, provide detailed information and supplementary explanations, and enhance the user's understanding depth.
[0225] Generate the final response: The response R after sentiment optimization and cognitive load optimization final As the final natural language response;
[0226] Through the above content, it is ensured that the responses output by the system are more humane and intelligent. In terms of emotional adaptation, it can provide soothing, positive or interactive responses according to the user's emotional state, enhance the user's emotional resonance and satisfaction. In terms of cognitive load adaptation, the complexity of the response content is dynamically adjusted, and the information is simplified to reduce the understanding pressure of users with high cognitive load, or detailed explanations are provided to meet the needs of users with low cognitive load. This dual optimization strategy makes the generated responses both in line with the user's emotions and easy to understand, which significantly improves the naturalness and adaptability of the user experience and enhances the system's interactive effects and service quality in complex scenarios.
[0227] like Figure 2As shown, the method for using the language model-based intelligent dialogue platform is implemented by the above-mentioned language model-based intelligent dialogue platform, and includes the following steps:
[0228] S1, user input reception and preprocessing: user input information is received by the user input processing unit, and the input information is preprocessed, including speech recognition and text segmentation, to convert the input information into a text vector;
[0229] S2, basic intention and potential intention analysis: The preprocessed user input information is input into the multi-stage dialogue intention analysis unit to analyze the semantic content of the input information, identify the user's basic intention, and combine historical dialogues, the user's emotional state and environmental perception data to infer the user's needs and hidden potential intentions;
[0230] S3, context management and adjustment: records the current conversation content in real time, and dynamically updates the context information based on the identified user needs and potential intentions;
[0231] S4, natural language response generation: Based on the identified user needs, potential intents, and adjusted context information, a preliminary natural language response is generated through natural language generation technology based on a language model;
[0232] S5, Emotion and cognitive load optimization: Combine the user's emotional state vector and cognitive load to optimize the tone and content of the generated preliminary natural language response, and finally generate a natural language response that meets the user's emotional state, comprehension ability and needs;
[0233] S6, output response: output the final optimized natural language response to the user, completing the intelligent interaction between the user and the platform.
[0234] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention. In order to make the public have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, but those skilled in the art can fully understand the present invention without the description of these details. In addition, in order to avoid unnecessary confusion about the essence of the present invention, well-known methods, processes, procedures, components and circuits are not described in detail.
[0235] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An intelligent dialogue platform based on a language model, characterized by: It includes a user input processing unit, a multi-stage dialogue intention analysis unit, a context management and tracking unit, and a response generation unit, wherein; The user input processing unit receives the user's input information and pre-processes the input information, including speech recognition and text segmentation; The multi-stage conversation intention analysis unit analyzes the semantic information of the user input information, identifies the user's basic intention, and identifies the user's needs and potential intentions in combination with the user's historical conversations, the user's emotional state, and environmental perception data, specifically including: Basic intent recognition: Analyze user input information through a convolutional neural network model to identify the user's basic intent, including queries, requests, and commands; Contextual reasoning: Based on the identified basic intent and the user's historical conversations, the real needs of the user can be identified; Potential intention recognition: Combine the user's emotional state and environmental perception data to infer the user's hidden potential intention; The context management and tracking unit records and manages the conversation context information in real time, and adjusts the context information according to the identified user needs and potential intentions; The response generation unit generates a corresponding natural language response based on the recognition results of user needs and potential intentions, combined with the adjusted context information, using a natural language generation technology based on a language model, and optimizes the response content according to the user's emotional state and cognitive load.
2. The language model-based intelligent dialogue platform according to claim 1, characterized in that: The user input processing unit comprises: Information reception: receiving user input information, including voice input information and text input information; Speech recognition: When receiving speech input information, a deep neural network model is used to convert the speech signal into a text vector; Text segmentation: When receiving text input information, the BERT model-based word segmentation technology is used to generate a text vector for each word or word fragment.
3. The language model-based intelligent dialogue platform according to claim 2, characterized in that: The convolutional neural network model includes: Input feature preparation: The text vectors generated by speech recognition and text segmentation are used as input features; Convolution operation: Use multiple convolution kernels of different sizes to convolve the input features to capture semantic information of different sizes; Non-linear transformation: The output after the convolution operation is non-linearly transformed through the ReLU activation function; Pooling operation: Use the maximum pooling method to reduce the dimension of the convolution features after nonlinear transformation; Intent classification: The feature map after the pooling operation is flattened and input into the fully connected layer to generate a prediction score for each intent category. The prediction score is converted into a probability distribution by applying the Softmax activation function. Intent recognition and output: According to the probability distribution of Softmax output, the category with the highest probability is selected as the final intent recognition result.
4. The language model-based intelligent dialogue platform according to claim 3, characterized in that: The contextual reasoning includes: Intention reasoning and context association: By combining the user’s basic intention and historical conversation content, a graph neural network model is used to obtain the context vector C. context ; Reasoning about real needs: Based on the obtained context vector C context Reasoning about the real needs of users real .
5. The language model-based intelligent dialogue platform according to claim 4, characterized in that: The potential intent identification includes: Emotional state recognition: Identify the emotional state vector E in the generated text vector emotion , including happiness, sadness, and anger; Environmental perception data processing: obtain the user's current environmental data and represent it as an environmental feature vector E env ; Potential intention inference: Combined with the user's emotional state vector E emotion and the environmental feature vector E env , the user's potential intention is inferred through the multimodal fusion model, and the weighted attention mechanism is used to generate the final potential intention vector R latent ; Output potential intent: potential intent vector R latent After classification by the Softmax classifier, the user’s hidden potential intention y is obtained latent .
6. The language model-based intelligent dialogue platform according to claim 5, characterized in that: The context management and tracking unit comprises: Context information recording: Real-time recording of the user's current conversation input information and the conversation platform's response results, building the current conversation context information C current ; Context information storage: Maintain a fixed-length context window W to store the most recent L rounds of conversation information; User needs and intention association update: Dynamically update and adjust the current context information based on the identified user needs and potential intentions; Context state optimization: Optimize the stored context information and extract the core content; Final context output: Based on the dynamically updated and optimized context information, the final context C is output final .
7. The language model-based intelligent dialogue platform according to claim 6, characterized in that: The response generation unit comprises: Basic response generation: Generates preliminary natural language responses based on the user's real needs, potential intentions, and final context; Emotion and cognitive load optimization: The tone and expression of the generated preliminary natural language response are optimized by combining the user's emotional state vector and cognitive load to generate the final natural language response.
8. The language model-based intelligent dialogue platform according to claim 7, characterized in that: The basic response generation includes: Input feature fusion: The user’s real needs R real , potential intention vector y latent and the final context C final Perform feature fusion to generate the response input vector X input ; Language model processing: The fused input vector X input Input into the natural language generation model to generate a preliminary natural language response R base .
9. The language model-based intelligent dialogue platform according to claim 8, characterized in that: The affect and cognitive load optimization includes: Emotional state adaptation: based on the user's emotional state vector E emotion , adjust the initial natural language response R base Tone of voice and emotional expression; Cognitive load adaptation: Combined with the user's cognitive load vector C cogload , optimize the complexity and information content of the initial natural language response after the emotional state adaptation; Generate the final response: The response R after sentiment optimization and cognitive load optimization final As the final natural language response.
10. A method for using a language model-based intelligent dialogue platform, implemented by the language model-based intelligent dialogue platform according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1, user input reception and preprocessing: user input information is received by the user input processing unit, and the input information is preprocessed, including speech recognition and text segmentation, to convert the input information into a text vector; S2, basic intention and potential intention analysis: The preprocessed user input information is input into the multi-stage dialogue intention analysis unit to analyze the semantic content of the input information, identify the user's basic intention, and combine historical dialogues, the user's emotional state and environmental perception data to infer the user's needs and hidden potential intentions; S3, context management and adjustment: records the current conversation content in real time and dynamically updates the context information based on the identified user needs and potential intentions; S4, natural language response generation: Based on the identified user needs, potential intents, and adjusted context information, a preliminary natural language response is generated through natural language generation technology based on a language model; S5, Emotion and cognitive load optimization: Combine the user's emotional state vector and cognitive load to optimize the tone and content of the generated preliminary natural language response, and finally generate a natural language response that meets the user's emotional state, comprehension ability and needs; S6, output response: output the final optimized natural language response to the user, completing the intelligent interaction between the user and the platform.
Citation Information
Patent Citations
Feature-based intention recognition algorithm
CN116612477A
Vehicle-mounted sign language interaction method and device and computer readable storage medium
CN117826984A
Intelligent customer service system based on large language model
CN117972100A
Intelligent dialogue platform based on large language model
CN118551772A
Optimization method and system of intelligent customer service system, equipment and medium
CN119228386A
Cited By
Artificial intelligence investigation and deep interview method and system
CN121146065A