Artificial intelligence dialogue generation method based on natural language processing
By employing technologies such as multi-layer word embedding, bidirectional LSTM encoders, multi-head attention mechanisms, and dynamic temperature sampling, combined with an edge-cloud collaborative architecture and privacy protection, the system addresses the issues of unstable response, privacy leakage, and resource constraints in AI dialogue systems, achieving efficient, secure, and interpretable dialogue generation.
Patent Information
- Application Number
- CN202511045190.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing AI dialogue systems based on natural language processing have shortcomings in terms of unstable response quality, privacy risks, poor adaptability to multiple scenarios, difficulty in deployment in resource-constrained environments, and lack of interpretability in response decisions.
It employs multi-layer word embedding, bidirectional hierarchical LSTM context encoder, multi-head gating attention mechanism, dynamic temperature sampling strategy, edge-cloud collaborative computing architecture and multi-level privacy protection mechanism, combined with reinforcement learning to generate efficient, secure and interpretable dialogue responses.
It achieves scene-adaptive control of dialogue response, improves the accuracy and security of intent matching, reduces latency and resource requirements, provides decision-making transparency and privacy protection, and ensures generation quality.
Smart Images

Figure CN120930182A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence dialogue system technology, and more specifically, to an artificial intelligence dialogue generation method based on natural language processing. Background Technology
[0002] In recent years, AI dialogue systems based on Natural Language Processing (NLP) have been widely applied in scenarios such as intelligent customer service, virtual assistants, and medical consultations. These systems generate human-like responses by understanding user input text and combining it with the dialogue context. Their core technology relies on the semantic encoding and decoding capabilities of deep learning models (such as Transformer and LSTM). With the development of pre-trained language models (BERT, GPT series), the fluency of dialogue generation has significantly improved. However, in practical industrial applications, bottlenecks such as unstable response quality, privacy risks, and poor adaptability to various scenarios still exist. Especially in highly sensitive fields (such as finance and healthcare) and resource-constrained environments (such as mobile terminals and IoT devices), existing technologies struggle to balance accuracy, security, and real-time performance requirements. Current technological shortcomings include:
[0003] 1. Static generation strategy leads to poor scene adaptability.
[0004] Existing dialogue systems often employ sampling strategies with fixed temperature parameters (e.g., Temperature = 0.7), which fail to dynamically adjust response characteristics based on the dialogue scenario. For instance, factual needs such as weather inquiries require deterministic outputs (e.g., "Shanghai will be 25℃ tomorrow"), but the system may generate redundant descriptions (e.g., "probably around 25℃") due to the fixed temperature value; conversely, creative dialogues, which require diversity, output rigid responses.
[0005] 2. Weak privacy protection mechanisms
[0006] The mainstream approach only performs keyword replacement during the input stage (e.g., replacing phone numbers with...). <phone>While the system implemented these features, it lacked continuous protection during the feature encoding and generation stages. Attackers could reconstruct sensitive information by reverse-engineering the model's attention weights (e.g., through gradient leakage attacks). Tests showed that when a user entered "My ID number is XXXXXXXXXXXXXXXXXX", the existing system leaked the complete number in the response with a probability as high as approximately 41%.
[0007] 3. Centralized architecture causes latency and cost issues.
[0008] Traditional cloud deployment solutions require transmitting the original user text to the server for processing throughout the entire process, resulting in a latency of 3-5 seconds per conversation in weak network environments. Furthermore, Transformer models with 24 or more layers require more than 8GB of video memory, making local deployment on mobile devices impossible.
[0009] 4. Lack of interpretability in response decisions
[0010] The current system does not indicate confidence levels or the basis for its decisions when generating responses, leaving users unable to ascertain the reliability of the answers. For example, in medical consultations, responses to questions like "Does a headache require medical attention?" do not differentiate between high-confidence professional advice (such as "Seek immediate medical attention") and low-confidence guesses (such as "May have caught a cold"). Medical field tests show that approximately 32% of AI diagnostic suggestions lead to misdiagnosis by users due to the lack of confidence level annotations.
[0011] Therefore, an artificial intelligence dialogue generation method based on natural language processing is proposed to address the above problems. Summary of the Invention
[0012] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an artificial intelligence dialogue generation method based on natural language processing to solve the problems mentioned in the background art.
[0013] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence dialogue generation method based on natural language processing, comprising: receiving dialogue text input by a user and converting it into a high-dimensional text vector representation through a multi-layer word embedding mechanism, including three-dimensional feature fusion of word-level embedding, positional encoding, and dialogue turn identifiers; employing a bidirectional hierarchical LSTM context encoder with gated recurrent units to extract deep semantic features of historical dialogues, the encoder generating context feature vectors containing grammatical structure and dialogue logic through forward and backward propagation at time steps; dynamically assigning weights to the text vectors and context feature vectors based on a multi-head gated attention mechanism, calculating cross-modal semantic correlation through a learnable parameter matrix and generating a joint semantic representation; using a Transformer decoder architecture containing 12-24 layers of self-attention modules to perform multi-round iterative decoding of the joint semantic representation, each decoder layer integrating relative positional encoding and residual connection techniques to generate a candidate response probability distribution; applying a context-aware dynamic temperature sampling strategy to select the optimal response text from the probability distribution, the strategy adaptively adjusting the sampling randomness according to the complexity of the dialogue scenario.
[0014] Preferably, the bidirectional hierarchical LSTM context encoder comprises: a first-layer LSTM unit that processes the original word embedding sequence and outputs a grammatical-level feature vector; a second-layer LSTM unit that receives grammatical features and generates a dialogue behavior-level feature vector; each layer adopts a forward and backward dual-path processing structure; wherein the initial state of the backward LSTM is obtained by linear transformation of the final hidden state of the forward LSTM, and the dual-path output generates a temporally dependent context feature matrix through a gated fusion module.
[0015] Preferably, the multi-head gating attention mechanism is specifically implemented as follows: mapping the text vector and the context feature vector to a query matrix Q, a key matrix K, and a value matrix V, respectively; calculating the score matrix S = Softmax(QK^T / √d_k) for each attention head, where d_k is a vector dimension scaling factor; performing weighted summation on the score matrix S to obtain a preliminary fusion vector H = SV; and inputting the outputs of each attention head into a sigmoid gating unit after compression by a one-dimensional convolutional layer with a kernel of 1 to generate the final joint semantic representation.
[0016] Preferably, the mathematical model of the dynamic temperature sampling strategy is: temperature coefficient τ = 0.1 + 0.9·σ(W_τ·c + b_τ), where c is the context feature vector, W_τ and b_τ are trainable parameters, and σ is the Sigmoid function; the sampling probability distribution is as follows: Nonlinear calibration is performed, and when τ approaches 0.1, the highest probability word is selected with a bias towards determinism, while when τ approaches 1.0, the randomness of the original distribution is preserved.
[0017] Preferably, reinforcement learning optimization is performed simultaneously during the decoding process: a ternary reward function is constructed, comprising a coherence reward R_c, an intent matching reward R_i, and an information reward R_d. R_c is calculated using the BLEU algorithm to determine the similarity between the generated response and the reference text in the knowledge base. R_i is calculated by comparing the cosine of the angle between the user input and the intent vector of the generated response using a pre-trained intent classifier. R_d is calculated based on the information entropy of the response text. The Transformer decoder parameters are updated using a proximal policy optimization algorithm, and the weights of the reward function are dynamically adjusted according to the dialogue scenario.
[0018] Preferably, a multi-level privacy protection mechanism is set up: during the input stage, sensitive words are jointly detected by regular expressions and named entity recognition. When privacy content is detected, the differential privacy module is triggered to replace the original words with hash-desensitized anonymous identifiers and add privacy marker bits to the context vector; during the output stage, a low-temperature sampling strategy is started for responses containing privacy markers to suppress the generation probability of high-risk words that may leak privacy.
[0019] Preferably, an edge-cloud collaborative computing architecture is adopted: the user terminal device performs word embedding and primary context encoding, generates compressed feature vectors, and transmits them to the cloud after AES-256 encryption; the cloud server completes the construction of joint semantic representation and Transformer decoding, and performs localization adaptation processing before the response text is returned to the terminal, including dialect conversion, personalized pronoun replacement, and text segmentation optimization adapted to the terminal screen size.
[0020] Preferably, a structured metadata tag is generated when outputting the response: it includes four-dimensional data including dialogue behavior type encoding, sentiment polarity intensity value (0-1 continuous scale), confidence score, and interpretability analysis vector; wherein the interpretability analysis vector is generated by visualizing the decoder attention weight and annotating the location index of key historical dialogue segments that affect the response generation; the metadata drives the downstream dialogue management module to make multi-round dialogue strategy decisions, and automatically triggers the manual takeover mechanism when the confidence score is lower than 0.7.
[0021] The technical effects and advantages of this invention are as follows:
[0022] Compared with existing technologies, this invention achieves scene-adaptive control of response characteristics through a dynamic temperature sampling strategy. In fact query scenarios, a low-temperature coefficient is used to output a deterministic response, while in open dialogue scenarios, a high-temperature coefficient is switched to retain diversity, significantly improving the accuracy of intent matching. Combined with a gating attention mechanism, key features of historical dialogues are dynamically filtered, and semantic fusion effects are enhanced through multi-head projection and convolutional compression, effectively solving the problem of long-range dependencies. An edge-cloud collaborative architecture is deployed to migrate word embedding and primary encoding to terminal devices for execution, and feature compression and encrypted transmission reduce cloud load, improving real-time performance while ensuring the computing power of deep learning models. A two-stage privacy protection system is established for input and output, using regular expressions and named entity recognition to jointly detect sensitive words, blocking information leakage paths through hash desensitization and risk word filtering, and strengthening the protection loop by using privacy flags linked to a temperature suppression mechanism. Innovatively, four-dimensional metadata tags are generated, integrating dialogue behavior classification, sentiment analysis, confidence assessment, and interpretability analysis functions to provide decision-making basis for downstream systems. When the confidence level is insufficient, a manual takeover process is automatically triggered. The various technical modules form an organic and collaborative system, optimizing the generation quality at the algorithm layer, ensuring safe and efficient deployment at the system layer, and enhancing decision-making transparency at the application layer. Attached Figure Description
[0023] Figure 1 This is a system framework diagram of the present invention.
[0024] Figure 2 This is a flowchart of the process of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Example 1:
[0027] As attached Figure 1-2 As shown, (1) an artificial intelligence dialogue generation method based on natural language processing includes: receiving dialogue text input by a user and converting it into a high-dimensional text vector representation through a multi-layer word embedding mechanism, which includes three-dimensional feature fusion of word-level embedding, position encoding, and dialogue turn identifier; using a bidirectional hierarchical LSTM context encoder with gated recurrent units to extract deep semantic features of historical dialogues, which generates context feature vectors containing grammatical structure and dialogue logic through forward and backward propagation of time steps; dynamically assigning weights to text vectors and context feature vectors based on a multi-head gated attention mechanism, and using a learnable parameter matrix. The method calculates cross-modal semantic correlation and generates a joint semantic representation. A Transformer decoder architecture with 12-24 layers of self-attention modules is used to iteratively decode the joint semantic representation in multiple rounds. Each decoder layer integrates relative position encoding and residual connection techniques to generate a candidate response probability distribution. A context-aware dynamic temperature sampling strategy is applied to select the optimal response text from the probability distribution. This strategy adaptively adjusts the sampling randomness according to the complexity of the dialogue scenario. The core implementation process is as follows: The user-input natural language query (e.g., "Is Shanghai suitable for outdoor activities the day after tomorrow?") first undergoes triple feature processing through an embedding layer. A pre-trained language model generates 768-dimensional word vectors, superimposed with word order-based position encoding information, and embeds a numeric identifier representing the current dialogue round. Subsequently, a bidirectional hierarchical LSTM context encoder processes the historical dialogue: the first LSTM layer analyzes the grammatical structure of the sentences and outputs a 256-dimensional feature vector; the second LSTM layer parses the dialogue behavior logic and outputs a 128-dimensional feature vector. The initial state of the inverse processing unit is generated by mapping the final state of the forward processing unit through a fully connected layer. Next, the current query and historical context are fused through 12 parallel attention mechanisms: the relevance weights of the query vector and the context vector are calculated, and noisy features with weight values below 0.1 are filtered out after Softmax normalization. The fused 1024-dimensional semantic representation is input into a 24-layer Transformer decoder, each layer containing a location-aware module and a residual connection structure, progressively generating a candidate response probability distribution. Finally, an adaptive temperature sampler is used to output the response, with the temperature coefficient calculated in real time based on the context features—outputting a low temperature value close to 0.15 for factual questions such as weather queries (ensuring deterministic responses), and outputting a high temperature value of around 0.8 for creative generation questions (preserving diversity).
[0028] (2) The bidirectional hierarchical LSTM context encoder comprises: a first-layer LSTM unit that processes the original word embedding sequence and outputs a grammatical-level feature vector; and a second-layer LSTM unit that receives grammatical features and generates a dialogue behavior-level feature vector. Each layer adopts a forward and backward dual-path processing structure. The initial state of the backward LSTM is obtained by linear transformation of the final hidden state of the forward LSTM. The dual-path output generates a temporally dependent context feature matrix through a gated fusion module. The implementation of the bidirectional hierarchical LSTM is divided into two processing stages: In the grammatical analysis layer, the forward processing unit sequentially parses the word sequence ["Shanghai", "the day after tomorrow", "outdoor activities"] to generate a hidden state, and the backward processing unit reverses the parse to generate a complementary state. The two are concatenated to form a 256-dimensional grammatical feature vector. In the dialogue behavior layer, the forward unit continues to process the grammatical feature sequence, and the initial state of its backward unit is generated by a fully connected layer activated by a hyperbolic tangent function from the final state of the grammatical layer. The two outputs are integrated through a gating fusion mechanism: the Sigmoid function is used to generate weight coefficients, and the forward and backward features are weighted and mixed to form a 128-dimensional dialogue behavior feature vector. The features of all historical dialogue rounds are stacked in chronological order to form a feature matrix, and an exponential decay coefficient is added to the earlier dialogues (e.g., the weight of the dialogue 10 rounds ago is reduced to 30%).
[0029] (3) The multi-head gating attention mechanism is specifically implemented as follows: the text vector and the context feature vector are mapped to the query matrix Q, the key matrix K and the value matrix V respectively; the score matrix S = Softmax(QK^T / √d_k) of each attention head is calculated, where d_k is the vector dimension scaling factor; the score matrix S is weighted and summed separately to obtain the preliminary fusion vector H = SV; the output of each attention head is compressed by a one-dimensional convolutional layer with a kernel of 1 and then input into the Sigmoid gating unit to generate the final joint semantic representation. The multi-head gating attention mechanism is implemented through five steps: the first step is to project the current query vector and the historical context matrix into 12 feature subspaces respectively; the second step is to calculate the attention score matrix of each subspace and control the numerical stability by the scaling factor; the third step is to normalize the score matrix and weight the fusion features; the fourth step is to compress the 12 outputs into compact features by depthwise separable convolution; the fifth step is to control the information flow intensity by using the gating unit - the Sigmoid function is used to generate the gating value and selectively superimpose it with the compressed features. Finally, residual joins and layer normalization operations are performed to generate a joint representation vector containing global semantics. For example, when processing travel information, this mechanism establishes a strong association between "attraction opening hours" queries and "holiday" information in historical conversations.
[0030] (4) The mathematical model of the dynamic temperature sampling strategy is: temperature coefficient τ = 0.1 + 0.9·σ(W_τ·c + b_τ), where c is the context feature vector, W_τ and b_τ are trainable parameters, and σ is the Sigmoid function; the sampling probability distribution is as follows: Nonlinear calibration is performed. When τ approaches 0.1, the highest probability words are selected with a bias towards determinism; when τ approaches 1.0, the randomness of the original distribution is preserved. The dynamic temperature sampler comprises two key stages: In the temperature coefficient calculation stage, the context feature vector is input to the fully connected layer, compressed by the Sigmoid function, and linearly mapped to the range of 0.15 to 0.85; the distribution calibration stage performs a nonlinear transformation on the original word probabilities, with the exponent parameter fixed at 1.5, and then divided by the temperature coefficient for smoothing. The calibrated probability distribution undergoes Top-k sampling (retaining the 50 candidate words with the highest probabilities). In practical applications, this manifests as follows: when a user queries stock prices, a low temperature of 0.15 is automatically used (outputting the deterministic answer "Tencent stock price 352 HKD"), while a high temperature of 0.85 is used when writing stories (generating diverse descriptions such as "The sunset dyed the sea red like blood").
[0031] (5) Reinforcement learning optimization is performed synchronously during the decoding process: A ternary reward function is constructed, comprising a coherence reward R_c, an intent matching reward R_i, and an information reward R_d. R_c is calculated using the BLEU algorithm to determine the similarity between the generated response and the reference text in the knowledge base. R_i is calculated using a pre-trained intent classifier to compare the cosine of the angle between the user input and the generated response's intent vectors. R_d is calculated based on the information entropy of the response text. The Transformer decoder parameters are updated using a near-end policy optimization algorithm, and the reward function weights are dynamically adjusted according to the dialogue scenario. The reinforcement learning module is started asynchronously after the dialogue is generated: First, the three reward indicators are calculated—the coherence reward is determined by comparing the similarity between the generated response and the reference text in the knowledge base using the BLEU-4 algorithm; the intent matching reward is determined by calculating the cosine similarity between the user query and the generated response using a pre-trained intent encoder; and the information reward is calculated based on the information entropy of the response text. The three reward values are dynamically weighted (weight coefficients 0.4 / 0.4 / 0.2) to form a comprehensive reward signal. A near-end strategy optimization algorithm is used to update model parameters: batch updates are performed every 200 sets of dialogue data, and a policy change threshold of 0.2 is set to prevent parameter mutations. For example, in a customer service scenario, a high-quality response to a "return process" query can enhance the corresponding parameters of the decoder by 15%.
[0032] (6) Multi-level privacy protection mechanism: In the input stage, sensitive words are jointly detected by regular expressions and named entity recognition. When privacy content is detected, the differential privacy module is triggered to replace the original word with an anonymous identifier that has been hashed and desensitized, and a privacy flag bit is added to the context vector. In the output stage, a low-temperature sampling strategy is initiated for responses containing privacy flags to suppress the generation probability of high-risk words that may leak privacy. The privacy protection implementation includes dual-stage control of input and output: In the input stage, regular expression matching (such as 11-digit mobile phone number mode) and named entity recognition model (detecting names / addresses) are used for dual detection of sensitive words. The identified privacy content is hashed by the HMAC-SHA256 algorithm and an anonymized identifier is output (such as "PHONE:7D8A3F"). A dedicated privacy flag bit (1024th dimension) is set in the feature vector. This flag bit triggers triple protection in the output stage: the temperature coefficient is forced to drop below 0.05, a preset risk word list is filtered (such as 200 sensitive words such as "ID card" and "bank card"), and the response length is limited to no more than 20 characters. When a user asks "How do I change my payment password?", the system returns a de-identified response: "For account security operations, please proceed through..." <app>of <settings>Processing.
[0033] (7) Adopt an edge-cloud collaborative computing architecture: The user terminal device performs word embedding and primary context encoding, generates a compressed feature vector, and then transmits it to the cloud after AES-256 encryption; the cloud server completes the construction of the joint semantic representation and Transformer decoding, and performs localization adaptation processing before returning the response text to the terminal, including text segmentation optimization such as dialect conversion, personalized pronoun replacement, and terminal screen size adaptation. Among them, the implementation of the edge-cloud collaborative architecture includes three layers of processing: The terminal device performs primary processing - generates word embeddings and 256-dimensional primary features through a lightweight model, compresses them to 30% of the original volume by the Zstandard algorithm, and then transmits them using AES-256 encryption (with the key rotated daily). The cloud server completes the core calculation - reconstructs the feature matrix and then performs deep LSTM and Transformer decoding. Localization adaptation is performed before the response text is returned to the terminal: calls the dialect word library for vocabulary replacement (such as converting "we" to "we哋" for Guangdong users), and splits long texts according to the terminal screen size (at most 30 Chinese characters per screen). The network transmission uses the QUIC protocol to ensure the stability of the weak network environment, and automatically enables the feature compression mode when the delay exceeds 100 milliseconds.
[0034] (8) Generate structured metadata tags when outputting the response: including four-dimensional data of the dialogue behavior type encoding, emotional polarity intensity value (continuous scale of 0-1), confidence score, and interpretability analysis vector; among them, the interpretability analysis vector is generated by visualizing the decoder attention weights, and annotates the position index of the key historical dialogue segments that affect the response generation; the metadata drives the downstream dialogue management module to make multi-round dialogue strategy decisions, and automatically triggers the manual takeover mechanism when the confidence score is lower than 0.7. Among them, the metadata generation system includes four dimensions of annotations: the output type encoding of the dialogue behavior classifier (0 inquiry / 1 reply / 2 suggestion, etc.); the emotional analysis model calculates the emotional polarity value from -1 to 1; the confidence score is converted by the entropy value of the response probability distribution (1 means completely certain); the interpretability analysis module annotates the position of the key historical statements that affect the decision (such as the second sentence in the 3rd round of dialogue). The application mechanism is as follows: when the confidence in the medical consultation scenario is lower than 0.7, the system automatically transfers to the manual customer service and prompts "Connecting you to an expert"; when the emotional polarity value is greater than 0.8, a smiling face emoji is added; the interpretability vector drives the generation of a decision summary (for example, "It is recommended to have a reexamination based on the medical history you described 3 minutes ago"). All metadata is encapsulated in JSON format for downstream system calls.
[0035] Example 2: Multi-source data joint modeling scenario
[0036] Step 1: User input and triple feature embedding
[0037] When the user inputs a natural language query such as "What's the weather like in Hangzhou on the weekend?", the system first performs triple feature processing:
[0038] 1. Word embedding: Each word is converted into a 768-dimensional vector (e.g., "Hangzhou" → [0.24, -0.85,..., 0.17]).
[0039] 2. Position encoding: Add position identifiers according to the word order (the first word +0.1, the second word +0.2, and so on).
[0040] 3. Turn identifier: Embed the value of the current conversation turn (e.g., add the identifier 5 for the 5th conversation turn). Example of the final output fusion vector: ["Hangzhou" vector + position 0.1 + turn 5, "weekend" vector + position 0.2 + turn 5,...].
[0041] Step 2: Hierarchical context encoding (bidirectional LSTM)
[0042] Input the historical conversation into a two-layer encoder:
[0043] Syntax layer:
[0044] Forward processing: Analyze the word sequence sequentially (["yesterday", "Hangzhou", "sunny"] → output syntax feature A).
[0045] Reverse processing: Analyze in reverse order (["sunny", "Hangzhou", "yesterday"] → output syntax feature B).
[0046] Concatenate the results: Concatenate feature A and B into a 256-dimensional vector.
[0047] Semantic layer:
[0048] Using the output of the previous layer as input, analyze the weather trend forward and the geographical association in reverse.
[0049] Gated fusion: Calculate the weights using the Sigmoid function (e.g., weight of recent conversations 0.9, weight of early conversations 0.3).
[0050] Generate a 128-dimensional context matrix to store the features of the recent 10 conversation turns
[0051] Step 3: Gated attention fusion
[0052] 1. Multi-head projection:
[0053] The current query vector is split into 12 groups of sub-vectors (each group is 64-dimensional).
[0054] The historical context matrix is split into 12 groups.
[0055] 2. Attention calculation:
[0056] For each group, the match score between the query and the history is calculated (e.g., "Weekend" matches "Saturday" historically with a score of 0.92).
[0057] Softmax normalization: Matching degree is converted into weight (0.92 → weight 35%).
[0058] 3. Feature compression:
[0059] The 12 outputs are compressed into 384 dimensions through convolution.
[0060] 4. Gated Filtering:
[0061] A filter coefficient (0-1) is generated, and irrelevant features below 0.1 (such as conversations from 3 days ago) are suppressed.
[0062] Output a 1024-dimensional joint semantic vector
[0063] Step 4: Response Generation (Dynamic Temperature Sampling)
[0064] Joint vector input 24-layer Transformer decoder:
[0065] 1. Iterative decoding:
[0066] Each decoder layer performs the following: Based on a location-aware mechanism, the system prioritizes identifying temporal relationships (e.g., "weekend" before "weather"), and retains key features from previous layers through residual connections to prevent information loss. The final output candidate word probability distribution is: sunny (0.85), rainy (0.10), cloudy (0.05). This architecture ensures both the accuracy of temporal logic and the integrity of feature information.
[0067] 2. Dynamic sampling:
[0068] The system employs a dynamic sampling strategy, adjusting the temperature coefficient based on the query type: simple queries (such as weather / time) use a low temperature coefficient of 0.15 to enhance determinism, while open-ended questions (such as story creation) use a high temperature coefficient of 0.85 to preserve diversity. The output distribution is optimized through probability calibration (the original probability raised to the power of 1.5 and divided by the temperature coefficient); for example, the calibrated probability for "sunny" is increased from 0.85 to 0.98. Finally, Top-50 sampling is performed from the calibrated distribution to ensure that generation quality is maintained while controlling diversity.
[0069] Step 5: Reinforcement Learning Optimization (Asynchronous Execution)
[0070] Start optimization after generating the response:
[0071] 1. Reward Calculation:
[0072] Coherence reward: Compare the 4-gram repetition rate of the generated response with the standard answer in the knowledge base. (For example, if the response "sunny turning cloudy over the weekend" matches the knowledge base "sunny on Saturday, cloudy on Sunday" → reward 0.8).
[0073] Intent reward: Calculate the cosine similarity between user query and response in the intent space.
[0074] (Query "How's the weather?" vs. response "Partly cloudy" → similarity 0.9).
[0075] Information reward: Statistical response information entropy (e.g., "sunny" has a low entropy value, "local thunderstorms" has a high entropy value).
[0076] 2. Parameter update:
[0077] Updated in batches every 200 conversations.
[0078] The variation range of the control parameters shall not exceed ±20%.
[0079] Example: When the response "bring an umbrella" receives a high reward in a rainy scenario, the corresponding generation probability increases.
[0080] Step 6: Two-stage control for privacy protection
[0081] 1. Input stage:
[0082] Sensitive word detection:
[0083] Regular expressions are used to match mobile phone numbers (11 digits) and ID card numbers (18 digits including X).
[0084] NER models can identify people's names / place names (such as "Zhang San" or "Chaoyang District").
[0085] Desensitization treatment:
[0086] The detection word is hashed using HMAC-SHA256 (input word + random salt value) → output
[0087] "NAME:7A3F".
[0088] 2. Output stage:
[0089] When privacy protection mode is activated, the system automatically implements strict security protocols: extremely low temperature sampling (temperature = 0.05) ensures deterministic responses, 200 sensitive words (such as financial and identity-related terms) are filtered in real time, and response length is limited to no more than 20 characters. For example, when a user inputs information involving sensitive operations, the system will return a standardized security prompt (such as "Account operations please pass through..."). <app>It is handled by the security center, while ensuring the function guide, completely avoiding the risk of privacy data leakage. This mechanism achieves zero exposure of privacy data through triple protection (low-temperature control, vocabulary filtering, length limitation).
[0090] Step 7: Edge-cloud collaborative processing
[0091] 1. Edge side (mobile phone / smart speaker):
[0092] Execute word embedding + primary LSTM.
[0093] Compress features to 30% of the original size (Zstandard algorithm).
[0094] AES-256 encryption (key is changed daily).
[0095] 2. Cloud side (GPU server):
[0096] Run the deep model after decompression.
[0097] Execute localization adaptation before returning the response:
[0098] Dialect conversion: Call the vocabulary list (e.g., for Sichuan users, "it's raining" → "raining cats and dogs").
[0099] Screen adaptation: Automatically segment (8 Chinese characters per line for smart watches, 20 Chinese characters for mobile phones).
[0100] 3. Network optimization:
[0101] Start the streamlined mode in a weak network environment (turn off the Transformer decoding above 12 layers).
[0102] Step 8: Metadata annotation and decision-making
[0103] Generate four types of metadata:
[0104] 1. Dialogue type: Classification model output encoding (0 inquiry / 1 reply / 2 suggestion...)
[0105] 2. Sentiment polarity: Sentiment analysis value [-1, 1] (-1 angry, 0 neutral, 1 happy)
[0106] 3. Confidence level: Converted based on the probability distribution entropy value (0.0 - 1.0)
[0107] 4. Explanation vector: Record the position of historical statements that affect the decision (e.g., the 2nd sentence in the 3rd round)
[0108] Application logic:
[0109] When the confidence level of medical consultation < 0.7 → Transfer to人工 and prompt "Contacting the doctor".
[0110] Affect value > 0.8 → Add emoticons (e.g., "Sunny") ”).
[0111] Interpretable vector-driven summary generation: "Based on your statement 'has a history of allergies' 2 minutes ago, it is recommended to avoid pollen."
[0112] Finally, the following points should be noted: First, in the description of this application, it should be noted that, unless otherwise specified and limited, the terms "installation", "connection", and "linkage" should be interpreted broadly, and can refer to mechanical or electrical connections, or internal connections between two components, or direct connections. "Up", "down", "left", "right", etc., are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may change.
[0113] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.
[0114] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.< / app> < / settings> < / app> < / phone>
Claims
1. A method for generating artificial intelligence dialogues based on natural language processing, characterized in that... include: The system receives user-input dialogue text and converts it into a high-dimensional text vector representation through a multi-layer word embedding mechanism. This representation includes a 3D feature fusion of word-level embedding, positional encoding, and dialogue turn identifiers. A bidirectional hierarchical LSTM context encoder with gated recurrent units is used to extract deep semantic features from the historical dialogue. This encoder generates context feature vectors containing grammatical structure and dialogue logic through forward and backward propagation at time steps. Based on a multi-head gated attention mechanism, dynamic weights are assigned to the text vectors and context feature vectors. Cross-modal semantic correlation is calculated using a learnable parameter matrix to generate a joint semantic representation. A Transformer decoder architecture containing 12-24 layers of self-attention modules is used to perform multi-round iterative decoding of the joint semantic representation. Each decoder layer integrates relative positional encoding and residual connection techniques to generate a candidate response probability distribution. A context-aware dynamic temperature sampling strategy is applied to select the optimal response text from the probability distribution. This strategy adaptively adjusts the sampling randomness according to the complexity of the dialogue scenario.
2. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... The bidirectional hierarchical LSTM context encoder comprises: a first-layer LSTM unit that processes the original word embedding sequence and outputs a grammatical-level feature vector; and a second-layer LSTM unit that receives grammatical features and generates a dialogue behavior-level feature vector. Each layer adopts a forward and backward dual-path processing structure. The initial state of the backward LSTM is obtained by linear transformation of the final hidden state of the forward LSTM. The dual-path outputs generate a temporally dependent context feature matrix through a gated fusion module.
3. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... The multi-head gating attention mechanism is specifically implemented as follows: the text vector and the context feature vector are mapped to the query matrix Q, the key matrix K, and the value matrix V, respectively; the score matrix S = Softmax(QK^T / √d_k) for each attention head is calculated, where d_k is the vector dimension scaling factor; the score matrix S is weighted and summed separately to obtain the preliminary fusion vector H = SV; the outputs of each attention head are compressed by a one-dimensional convolutional layer with a kernel of 1 and then input into the Sigmoid gating unit to generate the final joint semantic representation.
4. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... The mathematical model for the dynamic temperature sampling strategy is: temperature coefficient τ = 0.1 + 0.9·σ(W_τ·c + b_τ), where c is the context feature vector, W_τ and b_τ are trainable parameters, and σ is the Sigmoid function; the sampling probability distribution follows... Nonlinear calibration is performed, and when τ approaches 0.1, the highest probability word is selected with a bias towards determinism, while when τ approaches 1.0, the randomness of the original distribution is preserved.
5. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... Reinforcement learning optimization is performed simultaneously during the decoding process: a ternary reward function is constructed, which includes a coherence reward R_c, an intent matching reward R_i, and an information reward R_d. R_c is calculated using the BLEU algorithm to determine the similarity between the generated response and the reference text in the knowledge base. R_i is calculated by comparing the cosine of the angle between the user input and the intent vector of the generated response using a pre-trained intent classifier. R_d is calculated based on the information entropy of the response text. The Transformer decoder parameters are updated using a proximal policy optimization algorithm, and the weights of the reward function are dynamically adjusted according to the dialogue scenario.
6. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... A multi-level privacy protection mechanism is set up: during the input stage, sensitive words are jointly detected by regular expressions and named entity recognition. When privacy content is detected, the differential privacy module is triggered to replace the original words with hash-desensitized anonymous identifiers and add privacy flag bits to the context vector. During the output phase, a low-temperature sampling strategy is initiated for responses containing privacy markers to suppress the generation probability of high-risk words that may leak privacy.
7. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... An edge-cloud collaborative computing architecture is adopted: the user terminal device performs word embedding and primary context encoding, generates compressed feature vectors, and transmits them to the cloud after AES-256 encryption; the cloud server completes the construction of joint semantic representation and Transformer decoding, and performs localization adaptation processing before the response text is returned to the terminal, including dialect conversion, personalized pronoun replacement, and text segmentation optimization adapted to the terminal screen size.
8. The artificial intelligence dialogue generation method based on natural language processing as described in claim 1, characterized in that... When outputting the response, generate structured metadata tags: including dialogue behavior type encoding, sentiment polarity intensity value (0-1 continuous scale), confidence score, and four-dimensional data of interpretability analysis vector; The interpretability analysis vector is generated by visualizing the decoder attention weights and annotating the location index of key historical dialogue segments that affect response generation; the metadata drives the downstream dialogue management module to make multi-round dialogue strategy decisions, and automatically triggers a manual takeover mechanism when the confidence score is lower than 0.7.
Citation Information
Cited By
Rice blast incubation period diagnosis and prediction method and model based on double-flow data fusion
CN121582748A