Digital intelligent carrier pigeon WeChat intelligent dialogue generation method based on deep learning
The data obtained through deep learning technology and WeChat open platform API, combined with the public corpus and user portrait features, build a tree-like structure and introduce three-dimensional timestamps, and use a dynamic attention mechanism to generate dialogues, solving the problems of context coherence and personalized generation in multiple rounds of dialogues, achieving efficient and personalized dialogue generation effects.
Patent Information
- Application Number
- CN202510286072.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing dialogue generation model is difficult to maintain context coherence in multiple rounds of conversations, and it is difficult to personalize the conversation that meets user needs.
Through the deep learning-based digital intelligence televisor WeChat intelligent dialogue generation method, the WeChat open platform API is used to obtain authorized dialogue data, combine the public corpus for data cleaning and organization, build a tree structure and introduce three-dimensional timestamps to process multiple rounds of dialogue, and adopt dynamic attention mechanism and user portrait features to generate diverse and personalized dialogues.
It effectively improves the context coherence and personalization effect of the dialogue generation model, ensures the diversity and security of the generated dialogue, and solves the technical difficulties of context coherence and personalization generation in the existing technology.
Smart Images

Figure CN120216641A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of dialogue generation. More specifically, it relates to a digital intelligent pigeon WeChat intelligent dialogue generation method based on deep learning. Background Art
[0002] With the progress of artificial intelligence technology, especially the continuous evolution of deep learning models in the field of natural language processing (NLP), the application of intelligent dialogue systems has gradually developed from simple question-and-answer systems to complex multi-turn dialogue generation systems. Today's dialogue generation technology is widely used in fields such as intelligent customer service, virtual assistants, and social robots. In order to achieve more natural and intelligent dialogue generation, both academia and industry are constantly promoting the accuracy and efficiency of dialogue generation models.
[0003] Although deep learning technology, especially pre-trained models based on the Transformer architecture (such as GPT, BERT, etc.) has made significant progress, current dialogue generation models still face some challenges that need to be addressed urgently. The main problems include:
[0004] Context coherence problem: In multi-turn dialogues, how to maintain the coherence of the dialogue context. Especially when there are many dialogue turns, the transmission and memory of information are prone to loss, and the model may ignore previous dialogue content, resulting in irrelevant or out-of-context answers.
[0005] Personalization problem: As the dialogue progresses, the personalized needs of users (such as emotions, preferences, styles) become increasingly important. How to generate answers that meet their personalized needs based on user portraits and historical dialogues has become a difficult problem. Summary of the Invention
[0006] The present invention provides a digital intelligent pigeon WeChat intelligent dialogue generation method based on deep learning, aiming to solve the current context coherence problem and personalization problem.
[0007] The digital intelligent pigeon WeChat intelligent dialogue generation method based on deep learning includes the following steps:
[0008] Step 1: Obtain authorized dialogue data based on the WeChat Open Platform API, and use the public dialogue corpus as a supplementary data source; filter advertisements and spam information in the dialogue data and supplementary data source based on regular expressions; and perform semantic anomaly detection through a pre-trained BERT model to identify and filter out meaningless or abnormal dialogue fragments, obtaining a cleaned dialogue dataset;
[0009] Step 2: Based on the cleaned dialogue dataset, construct a tree structure to organize multi-turn dialogues. Each dialogue is organized into a tree structure according to timestamps and conversation context; and process the time information and dialogue order by attaching three-dimensional timestamps to each message to obtain the reconstructed multi-turn dialogue dataset;
[0010] Step 3: Construct corresponding sentiment values and text representations for the emojis in the multi-turn dialogue dataset. Add emotional information to the emojis by analyzing their sentiment values. Combine the text embeddings and visual features of the emojis based on a hybrid encoder, and enhance the semantic representation of the emojis by weighting the text embeddings and visual features to obtain the multi-turn dialogue dataset after emoji processing;
[0011] Step 4: Based on the multi-turn dialogue dataset after emoji processing, use a dynamic attention mechanism to encode each historical statement, calculate the influence of each historical statement, weight the historical statements through an importance scoring model, and adjust the memory length according to the dialogue turns to obtain the encoded context data;
[0012] Step 5: Based on the encoded context data, introduce user profile features, map the user profile features to the same vector space as the dialogue encoding, and map the user features to a representation compatible with the dialogue context through linear transformation to obtain the dialogue dataset integrated with user profiles;
[0013] Step 6: Based on the dialogue dataset integrated with user profiles, use a Transformer-based dialogue generation model to generate dialogue candidate sequences. Based on the generated dialogue candidate sequences, use the Beam Search algorithm, introduce a diversity penalty term in the frequency division of each candidate sequence, and dynamically adjust the generation parameters according to the user profile to generate diverse dialogue sequences;
[0014] Step 7: Filter harmful content from the generated dialogue sequences through a rule layer, a model layer, and a post-processing layer to generate safe dialogue sequences;
[0015] Step 8: Based on the safe dialogue sequences, construct a composite loss function based on cross-entropy loss, coherence loss, and style loss to optimize the dialogue generation model to obtain the optimized dialogue generation model; Generate dialogues based on the optimized dialogue generation model.
[0016] The present invention obtains authorized conversation data through the WeChat Open Platform API, combines it with a publicly available corpus, and performs semantic anomaly detection through regular expressions and the BERT model to ensure that the cleaned conversation dataset does not contain advertisements, spam, or meaningless conversation fragments, thus guaranteeing the high quality and relevance of the data. Secondly, when processing multi-turn conversations, the conversation data is organized by constructing a tree structure and introducing a three-dimensional timestamp, ensuring the effective processing of the conversation order and time information, thereby enhancing the coherence of the context. Regarding the emotional information of emojis, the text and visual features of the emojis are combined through a hybrid encoder, enhancing the emotional expression in the conversation and further improving the effect of personalized conversation generation. To better model the conversation context, a dynamic attention mechanism is adopted, which can weight according to the importance of historical statements and adjust the memory length, thereby effectively retaining context information. After introducing user profile features, by fusing user features with conversation encoding, the generated conversations are made more personalized to meet the needs of different users. In the generation stage, through a Transformer-based model and the Beam Search algorithm, a diversity penalty term is introduced during the generation process and the generation parameters are dynamically adjusted according to the user profile, thus ensuring the diversity and security of the generated conversations. Finally, based on the optimization mechanism of the composite loss function, the accuracy and coherence of the conversation generation model are further improved, effectively solving the technical problems of context coherence and personalized generation.
[0017] Preferably, the generation steps of the reconstructed multi-turn conversation dataset are as follows:
[0018] Temporal encoding: Append a three-dimensional timestamp to each message, including an absolute timestamp, a relative time difference, and an in-conversation order.
[0019] Temporal embedding fusion: Based on the pre-trained BERT model, perform text embedding on each message to obtain a text representation, and then perform embedding on the timestamp, relative time difference, and in-conversation order to obtain a temporal representation; weight and fuse the temporal representation and the text embedding to obtain the reconstructed multi-turn conversation dataset.
[0020] Preferably, step 3 includes the following steps:
[0021] Step 3.1: Construct a mapping table containing emojis and emotional values, and define emotional values for each emoji.
[0022] Step 3.2: Based on the pre-trained BERT model, convert the text description of the emoji into a high-dimensional vector representation to obtain a text embedding vector.
[0023] Step 3.3: Process the image of the emoji using a pre-selected ResNet to obtain a visual feature vector.
[0024] Step 3.4: The text embedding vector and the visual feature vector are fused in a weighted average manner to obtain an expression conformity representation vector; among them, the weights in the weighted average are dynamically adjusted based on the dialogue context;
[0025] Step 3.5: Use a bidirectional LSTM model to correct the emotional value of the expression conformity through context information to obtain an adjusted expression conformity representation vector, and then obtain a multi-turn dialogue dataset after expression conformity processing.
[0026] Preferably, the said step 4 includes the following steps:
[0027] Step 4.1: Adopt a sliding window mechanism to retain a predetermined number of dialogue turns;
[0028] Step 4.2: In the determined dialogue turns, based on the vector representation of each dialogue, use an importance scoring model to calculate the attention weights:
[0029]
[0030] In the formula: α i represents the attention weight of the historical dialogue h i ; W q represents the query weight matrix; W k represents the key weight matrix; b represents the bias term; σ represents the Sigmoid activation function; R i represents the importance score of the historical dialogue h i ; The importance score is calculated based on time factors, emotional factors, and semantic relevance;
[0031] Step 4.3: Use the calculated attention weights to weight the corresponding historical dialogue to obtain a weighted dialogue representation, that is, the encoded context data.
[0032] Preferably, the said step 5 includes the following steps:
[0033] Step 5.1: Extract user features, including interest tags, behavior data, preference information, and social attributes, and vectorize the extracted user features;
[0034] Step 5.2: Perform a weighted average on the user's interest tags, behavior data, preference information, and social attribute vectors to obtain a comprehensive user portrait vector;
[0035] Step 5.3: Use a linear transformation to map the user portrait to a vector space that is the same as the encoded context data to obtain a projected user portrait vector;
[0036] Step 5.4: The encoded context data and the projected user profile are fused in a weighted fusion manner; a dialogue dataset with a fused user profile is obtained, where the weights during weighted fusion are dynamically adjusted based on a neural network of the dialogue content and user characteristics.
[0037] Preferably, the specific steps of using the Beam Search algorithm are as follows:
[0038] Generate candidate words: Based on the dialogue candidate sequence generated by the Transformer-based dialogue generation model, obtain the generation probability of each candidate word;
[0039] Diversity penalty: Calculate the diversity penalty for each candidate word:
[0040]
[0041] In the formula: represents the hidden state h of the current generated word y t and the maximum pre-similarity between the hidden state h of the historical generated word y t ; λ represents the hyperparameter of the diversity penalty; j and the hidden state h of the historical generated word y j ; λ represents the hyperparameter of the diversity penalty;
[0042] Temperature adjustment: The generation probability of each candidate word is adjusted according to the user's style requirements. By adjusting the generation probability distribution, the randomness and diversity of the generation process are made to conform to the user's preferences:
[0043]
[0044] In the formula: T u represents the temperature adjustment factor;
[0045] Comprehensive scoring and ranking: The comprehensive score of each candidate sequence is adjusted based on the generation probability, historical similarity, and style requirements:
[0046]
[0047] In the formula: α represents the hyperparameter of the style adjustment factor; logP(y t |y1,y2…y t-1 ) represents the logarithmic probability when the candidate word y t is given the historical word sequence y1,y2…y t-1 ;
[0048] Based on the above steps, it is repeated at each time step. A new candidate sequence is extended according to the candidate sequence of the previous moment, and sorted according to the score. Finally, the optimal candidate sequence is selected.
[0049] Preferably, the rule layer is a fast screening based on keyword and pattern matching, and uses regular expressions and keyword matching to quickly detect potential harmful content;
[0050] The model layer is a classification model based on the Transformer architecture, which classifies the content of each generated dialogue to determine whether the text contains inappropriate content;
[0051] The post-processing layer uses a dialogue understanding model based on BERT to analyze the rationality of the dialogue context, ensuring that the generated dialogue does not violate the previous dialogue background at the semantic level; then, through the review and analysis of the historical dialogue, it judges whether the generated content is reasonable, and if it is not reasonable, manual intervention is carried out.
[0052] Preferably, the composite loss function is as follows:
[0053]
[0054] In the formula: λ CE 、λ coh 、λ style respectively represent the weight hyperparameters of cross-entropy loss, coherence loss, and style loss; represents the cross-entropy loss; represents the coherence loss; represents the style loss;
[0055]
[0056] In the formula: y i represents the true label distribution; represents the probability distribution predicted by the model; N represents the size of the vocabulary;
[0057] Among them, the coherence loss is measured by calculating the semantic distance between the currently generated sentence and the previous round of dialogue:
[0058]
[0059] In the formula: e(s t ) represents the semantic embedding representation of sentence s t ; e(s t-1 ) represents the semantic embedding representation of sentence s t-1 ; CosineSimilarity represents the cosine similarity function;
[0060] Among them, the style loss is calculated based on the embedding vectors of user portrait features and sentiment analysis:
[0061]
[0062] In the formula: e(ustyle ) represents the embedding of style-related features in the user profile; Represents the square of the Euclidean distance.
[0063] The beneficial effects of the present invention include:
[0064] The present invention obtains authorized dialogue data through the WeChat Open Platform API, combines it with a public corpus, and through semantic anomaly detection using regular expressions and the BERT model, ensures that the cleaned dialogue dataset does not contain advertisements, spam, or meaningless dialogue fragments, thus guaranteeing the high quality and relevance of the data; secondly, when processing multi-turn conversations, the dialogue data is organized by constructing a tree structure and introducing a three-dimensional timestamp, ensuring the effective processing of the dialogue order and time information, thereby enhancing the coherence of the context; for the emotional information of emojis, the text and visual features of the emojis are combined through a hybrid encoder, enhancing the emotional expression in the dialogue and further improving the effect of personalized dialogue generation; in order to better model the dialogue context, a dynamic attention mechanism is adopted, which can weight according to the importance of historical sentences and adjust the memory length, thereby effectively retaining context information; after introducing user profile features, by fusing user features with dialogue encoding, the generated dialogue becomes more personalized, meeting the needs of different users; in the generation stage, through a Transformer-based model and the Beam Search algorithm, a diversity penalty term is introduced during the generation process and the generation parameters are dynamically adjusted according to the user profile, thus ensuring the diversity and security of the generated dialogue; finally, based on the optimization mechanism of the composite loss function, the accuracy and coherence of the dialogue generation model are further improved, effectively solving the technical problems of context coherence and personalized generation. Description of the Drawings
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1 It is the overall step block diagram provided by the embodiment of the present invention.
[0067] Figure 2 It is the specific step block diagram of step 3 provided by the embodiment of the present invention.
[0068] Figure 3 It is the specific step block diagram of step 5 provided by the embodiment of the present invention.
[0069] Figure 4Schematic diagram of the encoder structure in step 6 provided by an embodiment of the present invention.
[0070] Figure 5 Schematic diagram of the decoder structure in step 6 provided by an embodiment of the present invention. Detailed implementation manners
[0071] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0072] Refer to Figure 1 As shown, the dynamic risk pricing optimization method for insurance products with multi-factor integration includes the following steps:
[0073] The intelligent dialogue generation method for the digital intelligent pigeon on WeChat based on deep learning includes the following steps:
[0074] Step 1: Obtain authorized dialogue data based on the API of the WeChat Open Platform, and use the public dialogue corpus as a supplementary data source; filter out advertisements and spam information in the dialogue data and the supplementary data source based on regular expressions; and perform semantic anomaly detection through a pre-trained BERT model to identify and filter out meaningless or abnormal dialogue fragments to obtain a cleaned dialogue data set;
[0075] By calling the API of the WeChat Open Platform, obtain the authorized dialogue data between the user and the WeChat intelligent dialogue system; the dialogue data includes the text input by the user and the response messages of the system, forming a complete dialogue record; the dialogue data contains multiple rounds of interactions, covering various scenarios and topics, and is used as the training data of the dialogue generation model; on this basis, in order to enrich the data source, use the public dialogue corpus as a supplementary data source (such as Weibo, etc.);
[0076] After obtaining the dialogue data and the supplementary corpus, filter out the advertisements and spam information therein, which is achieved through regular expressions, and identify and filter out irrelevant advertisement content or meaningless and repetitive information messages by setting keyword patterns (such as "preferential", "free", "promotion", "click the link", etc.);
[0077] For example: the system matches patterns such as "click this link to get a discount" or "register immediately" through regular expressions to automatically delete the dialogue content with advertisement nature;
[0078] After filtering out the advertisements and spam information, perform semantic anomaly detection. Adopt a pre-trained BERT model to identify and filter out semantically incoherent or abnormal dialogue fragments by analyzing the semantic structure in the dialogue text;
[0079] For example, if a grammar error, spelling mistake is found in the dialogue, or there are serious conflicts in the sentence logic of the dialogue, it will be marked as abnormal content and filtered;
[0080] After the above processing, a cleaned dialogue dataset is obtained. The dataset excludes irrelevant information, incorrect content, and data unsuitable for training, ensuring the quality of the data.
[0081] Step 2: Based on the cleaned dialogue dataset, construct a tree structure to organize multi-turn dialogues. Each dialogue is organized into a tree structure according to the timestamp and conversation context; and process the time information and dialogue order by attaching a three-dimensional timestamp to each message to obtain a reconstructed multi-turn dialogue dataset;
[0082] In the cleaned dialogue dataset, each dialogue contains:
[0083] Message ID: The unique identifier of each message;
[0084] Parent message ID: If the current message is a reply to a certain message, it will have a parent message ID to mark which message this message is a reply to;
[0085] Sender ID: Indicates the sender of the message;
[0086] Timestamp: The sending time of the message;
[0087] Message content: The actual text content;
[0088] The tree structure of the dialogue is constructed in the following way:
[0089] Establish the association between parent and child messages through the relationship between the message ID and the parent message ID to form a tree structure. The parent message ID of each message will point to the parent node of the message;
[0090] The root node of the tree structure is the first message of the dialogue, and each subsequent message is constructed according to the parent message ID in turn;
[0091] Attach a three-dimensional timestamp to each message: absolute timestamp (recording the sending time of each message), relative time difference (recording the time difference between each message and the previous message), in-conversation order (recording the order of each message in the dialogue);
[0092] Exemplarily, the three-dimensional timestamp is as follows:
[0093] Message 1: Absolute timestamp = 1589734200, relative time difference = N / A, in-conversation order = 1
[0094] Message 2: Absolute timestamp = 1589734290, relative time difference = 90 seconds, in-conversation order = 2
[0095] Message 3: Absolute timestamp = 1589734500, relative time difference = 210 seconds, in-conversation sequence = 3;
[0096] Temporal encoding: For each message, calculate the three-dimensional timestamp, and use temporal encoding techniques (such as positional encoding or time encoding) to convert the three-dimensional timestamp information into an embedding vector;
[0097] Temporal embedding fusion: Use a pre-trained BERT model to perform text embedding on each message to obtain a text representation; convert the three-dimensional timestamp into an embedding vector through a temporal embedding process. Specifically, use a time encoding method based on sine and cosine functions to encode the time features;
[0098] Perform weighted fusion of the text embedding and the temporal embedding. During the weighted fusion process, the weights are set based on task requirements.
[0099] Step 3: Construct corresponding sentiment values and text representations for the emojis in the multi-turn dialogue dataset. Add sentiment information to the emojis by analyzing their sentiment values. Combine the text embedding and visual features of the emojis based on a hybrid encoder, and enhance the semantic representation of the emojis by weighting the text embedding and visual features to obtain the multi-turn dialogue dataset after emoji processing;
[0100] See Figure 2 As shown, Step 3 includes the following steps:
[0101] Step 3.1: Construct a mapping table that includes emojis and sentiment values, and define sentiment values for each emoji;
[0102] Among them, the mapping table is constructed based on an emoji dictionary and a corresponding sentiment value dictionary. Each emoji is mapped to a sentiment value, and the sentiment value can be numerical or categorical;
[0103] For example, a numerical value greater than 0 represents positive, a value less than 0 represents negative, and a value equal to 0 represents neutral;
[0104] That is, the numerical value ranges from [-1, 1]. The closer the value is to 1, the higher the positive emotion value; the closer the value is to -1, the higher the negative emotion value;
[0105] The categorical method is positive, negative, neutral, etc. This is only exemplary, and more categories can be defined to represent emotions; the mapping is constructed through manual annotation.
[0106] Step 3.2: Convert the text description of the emoji into a high-dimensional vector representation based on a pre-trained BERT model to obtain a text embedding vector;
[0107] Assume that each emoji has a text description, for example: smile, anger, etc.;
[0108] Given the text description of an emoji, it is processed by a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model to obtain a text embedding vector; where the input of the pre-trained BERT model is a sequence T e =[w1, w2, …, w n composed of words in a vocabulary, and the output of the pre-trained BERT model is the context embedding representation corresponding to each vocabulary; where the context embedding representation (i.e., the text embedding vector) of each vocabulary in the output is obtained through the hidden state of the last layer of the pre-trained BERT model.
[0109] Step 3.3: Process the image of the emoji using a pre-trained ResNet to obtain a visual feature vector;
[0110] In this embodiment, the image of the given emoji is processed by a pre-trained ResNet model (Residual Network model) to obtain a visual feature vector. The features in the image are extracted through convolutional layers and residual connections, and finally a high-dimensional vector representation is output, which is the visual feature vector.
[0111] Step 3.4: Fuse the text embedding vector and the visual feature vector in a weighted average manner to obtain an emoji representation vector; where the weights in the weighted average are dynamically adjusted based on the dialogue context;
[0112] As a possible implementation manner of this embodiment, the weights in the weighted average are adjusted through an adaptive gating weight process, and an adaptive weight g t and g v are calculated through a gating network to dynamically adjust the weights of the text embedding vector and the visual feature vector:
[0113] g t =σ(W t ·C + b t );
[0114] g v =σ(W v ·C + b v );
[0115] In the formula: C represents the dialogue context vector; b t and b v represent bias terms; W t and W vdenotes the weight matrix; σ denotes the Sigmoid activation function;
[0116] E e = g t T e + g v V e ;
[0117] where: T e denotes the text embedding vector; V e denotes the visual feature vector; E e denotes the fused feature vector;
[0118] Step 3.5: Use a bidirectional LSTM model to correct the emotion value of the emoji through context information, obtain the adjusted emoji representation vector, and further obtain the multi-turn dialogue dataset after emoji processing.
[0119] Encode the dialogue context and emoji representation vector through bidirectional LSTM:
[0120]
[0121] where, h t denotes the emoji representation vector after bidirectional LSTM encoding;
[0122] Correct the emotion value of the emoji through an adaptive gating mechanism:
[0123] g c = σ(W c ·h t + b c );
[0124]
[0125] where: W c denotes the adaptive gating weight matrix; b c denotes the bias term; σ denotes the Sigmoid activation function; g c denotes the adaptive weight; denotes the corrected emotion vector, i.e., the adjusted emoji representation vector; C denotes the context vector;
[0126] In this embodiment, since the emojis in modern dialogue systems (such as WeChat) contain not only text information but also image information (such as the visual images of emojis); these information come from different modalities and have different characteristics. If simple weighted averaging or linear combination is used, it is often unable to capture the importance of each modality in a specific context well. Therefore, in this embodiment, an adaptive gating mechanism is added to dynamically determine the contribution of each modality to the final emoji representation; and since the emotion of an emoji is not only its static feature but is also strongly influenced by the dialogue context. Even if the visual or text features of an emoji itself have a certain emotional color, in some dialogue situations, it needs to be adjusted to match the current context. And we use an emotion correction gating mechanism to combine context information to achieve dynamic correction of the emotion of emojis, thereby enhancing the semantic performance of emojis in conversations.
[0127] For example: For a smiling emoji, its text feature is "smile". When we only consider the visual and text features brought by the smiling emoji itself, we will ignore the application scenario, which may lead to incorrect judgments; after a smiling emoji appears, the context information may be negative or positive. When it is negative, the smiling emoji may express sarcasm and dissatisfaction. Therefore, we cannot completely consider the smiling emoji itself and need to combine the context for accurate judgment.
[0128] And it should be noted that in this embodiment, the gating mechanism is involved in both step 3.4 and step 3.5. The purpose of the gating mechanism in step 3.4 is to adjust the fusion of multi-modal features to ensure that the features of different modalities can be effectively fused according to the context; while the gating mechanism in step 3.5 mainly focuses on how to adjust the emotion of emojis to ensure that their emotional expressions in the conversation can be optimized as the conversation progresses; the two complement each other and enhance the accuracy of dialogue understanding at different levels.
[0129] Step 4: Based on the multi-round dialogue dataset processed by emojis, use a dynamic attention mechanism to encode each historical sentence, calculate the influence of each historical sentence, weight the historical sentences through an importance scoring model, and adjust the memory length according to the dialogue round to obtain the encoded context data;
[0130] The said step 4 includes the following steps:
[0131] Step 4.1: Use a sliding window mechanism to retain a predetermined number of dialogue rounds;
[0132] In this embodiment, a method for dynamically adjusting the dialogue round is provided, which dynamically adjusts the size of the window according to the time difference and relevance:
[0133] Time difference calculation: For each round of conversation, calculate the time difference from the previous round of conversation. If the time difference is large, it indicates that the influence of the historical conversation may weaken, and thus the window needs to be narrowed:
[0134] Δt i = t i - t i-1 ;
[0135] Where: t i and t i-1 respectively represent the timestamps of the i-th round of conversation and the (i - 1)-th round of conversation; Δt i represents the time difference between the i-th round of conversation and the (i - 1)-th round of conversation;
[0136] Content relevance calculation: Evaluate the semantic coherence between the current round of conversation and the previous round of conversation based on the relevance of the conversation content. If the relevance is high, expand the window range to retain more relevant historical conversations:
[0137] R i = Sim(h i , h i-1 );
[0138] Where: Sim(h i , h i-1 ) represents the text similarity between the i-th round of conversation and the (i - 1)-th round of conversation;
[0139] Sliding window range selection: Calculate the size of the time window based on the calculated time difference and content relevance:
[0140] N i = N base + α·Δt i + β·R i ;
[0141] Where: N i represents the window size at the current conversation round i; N base represents the base window size; α and β respectively represent the adjustment coefficients of the influence of time difference and relevance on the window size;
[0142] According to the window size N i , the sliding window mechanism selects the context that includes the current conversation and its previous N i rounds of conversations from the conversation history:
[0143]
[0144] Among them: Window i represents the range of conversation history selected by the system for the i-th round of conversation.
[0145] Step 4.2: In the determined dialogue turn, based on the vector representation of each dialogue, use the importance scoring model to calculate the attention weights:
[0146]
[0147] Where: α i represents the attention weight of the historical dialogue h i ; W q represents the query weight matrix; W k represents the key weight matrix; b represents the bias term; σ represents the Sigmoid activation function; R i represents the importance score of the historical dialogue h i ; The importance score is calculated based on time factors, emotional factors, and semantic relevance;
[0148] Exemplarily, the calculation method of the importance score is as follows:
[0149] R i = γ1·Sim(h i , h i-1 ) + γ2·f time (t i , t i-1 ) + γ3·f emotion (h i );
[0150] Where: Sim(h i , h i-1 ) represents the semantic similarity between the i-th round of dialogue and the (i - 1)-th round of dialogue, calculated by cosine similarity; h i and h i-1 represent the text embedding vectors of the i-th round of dialogue and the (i - 1)-th round of dialogue respectively; γ1, γ2, γ3 represent the importance adjustment coefficients of each dimension;
[0151]
[0152] Where: t i and t i-1 represent the timestamps of the i-th round of dialogue and the (i - 1)-th round of dialogue; α represents the adjustment parameter for controlling the rate of time decay; f time (t i , t i-1 ) represents the time decay function, measuring the impact of the time interval between the i-th round of dialogue and the (i - 1)-th round of dialogue on the importance of the historical dialogue;
[0153] f emotion (h i ) = λ·S emoji (h i );
[0154] Wherein: S emoji (h i ) represents the weighted sum of the sentiment values of all emojis in the i-th round of conversation; λ represents the sentiment weight;
[0155] Step 4.3: Weight the corresponding historical conversation using the calculated attention weights to obtain the weighted conversation representation, that is, the encoded context data; that is, use the attention weights α obtained in the above Step 4.2 i Weight the corresponding historical conversation to obtain the encoded context data.
[0156] Step 5: Based on the encoded context data, introduce user portrait features, map the user portrait features to the same vector space as the conversation encoding, and map the user features to a representation compatible with the conversation context through linear transformation to obtain a conversation dataset integrating the user portrait;
[0157] See Figure 3 As shown, Step 5 includes the following steps:
[0158] Step 5.1: Extract user features, including interest tags, behavior data, preference information, and social attributes, and vectorize the extracted user features;
[0159] Among them, the interest tags include the user's interest fields on the platform, such as music, movies, technology, etc.;
[0160] The behavior data includes the user's historical behavior data, such as click records, chat records, purchase records, etc.;
[0161] The preference information includes the user's preference settings, such as the preferred conversation style, language, etc.;
[0162] The social attributes include the user's social attribute information, such as age, gender, occupation, etc.;
[0163] For vectorizing the interest tags, we adopt a predefined tag set to convert the user's interest tags into a multi-hot encoding vector. Assuming there are N interest tags, the vector representation of the user's interest tags is
[0164] For vectorizing the behavior data, using the user's historical behavior data, convert it into a fixed-length vector through the Embedding method; assuming the vector representation obtained by Embedding the behavior data is
[0165] For vectorizing the preference information, convert the user's preference information into a multi-hot encoding vector, and the vector representation of the preference information is
[0166] The social attribute vector is converted using a multi-hot encoding vector and is denoted as
[0167] Step 5.2: Perform a weighted average on the user's interest tags, behavior data, preference information, and social attribute vector to obtain a comprehensive user profile vector;
[0168] v user = α·v interest + β·v behavior + γ·v preference + δ·v social ;
[0169] In the formula: α, β, γ, and δ all represent weight coefficients;
[0170] The weight coefficients can be preset manually or adjusted through a fully connected network;
[0171] Exemplarily, the feature data of each user is input into the fully connected network, where the feature data includes behavior data, interest tags, preferences, and social attributes;
[0172] The structure of the fully connected network includes multiple fully connected layers. The activation function uses the ReLU activation function. The input data is processed through the fully connected layers, and then four weight coefficients are output through the output layer, corresponding to the above α, β, γ, and δ respectively. And learning the change of weights based on the fully connected network to dynamically adjust the weights is a conventional technical means in the art, so it will not be elaborated here.
[0173] Step 5.3: Map the user profile to a vector space that is the same as the encoded context data using a linear transformation to obtain the projected user profile vector;
[0174] Let the dialogue context vector obtained in Step 4 be To make the user profile vector compatible with the dialogue context vector, the user profile vector is projected into a vector space that is the same as the context vector. Therefore, we use a linear transformation matrix to map the user profile:
[0175] v user_mapped = W user ·v user + b user ;
[0176] In the formula: W user represents the transformation matrix; b user represents the bias term;
[0177] As a further implementation manner of this embodiment, to introduce non-linear features, we can apply the activation function ReLU after the mapping operation; that is:
[0178] v user_mapped = ReLU(W user ·v user + b user );
[0179] At this time, v user_mapped After the above linear transformation and the introduction of non - linearity, the user portrait vector will be in the same vector space as the dialogue context vector.
[0180] Step 5.4: Fuse the encoded context data and the projected user portrait in a weighted fusion manner; obtain the dialogue dataset with the fused user portrait, where the weights during weighted fusion are dynamically adjusted by learning through a neural network based on the dialogue content and user characteristics;
[0181] In this embodiment, in order to fuse the encoded dialogue context vector v context and the projected user portrait vector v user_mapped a weighted fusion method is adopted to assign a dynamic weight to each input feature, where the weights are obtained by learning through a multi - layer perceptron;
[0182] Define the weighting coefficient indicating the fusion weight of the context and the user portrait:
[0183] w fusion = MLP(v context , v user_mapped );
[0184] where MLP represents a multi - layer perceptron; it receives the context vector and the user portrait vector as inputs and outputs a weight vector, including two elements, respectively representing the weights of the context and the user portrait;
[0185] Perform weighted fusion based on the obtained weights of the context and the user portrait:
[0186] v fused = w fusion [1]·v context + w fusion [2]·v user_mapped ;
[0187] In the formula: w fusion [1] represents the weighting coefficient of the context; w fusion [2] represents the weighting coefficient of the user portrait; v fused represents the dialogue dataset with the fused user portrait;
[0188] As a further implementation of this embodiment, in order to keep the influence of the weighting coefficients consistent, the weighting coefficients are normalized so that the sum of the two weighting coefficients is 1;
[0189] In this embodiment, the linear transformation in step 5.3 maps the user profile to the vector space of the dialogue context, and then through the dynamic weighted fusion mechanism in step 5.4, the user and the dialogue content are closely combined to generate more personalized dialogue content. By dynamically learning the weighting coefficients through the neural network, the fusion of the user profile and the dialogue context is more flexible and personalized, and the contribution degree of each feature can be adjusted according to the actual situation, thereby improving the quality of dialogue generation and user satisfaction.
[0190] Step 6: Use a Transformer-based dialogue generation model to generate a dialogue candidate sequence based on the dialogue dataset that fuses the user profile. Based on the generated dialogue candidate sequence, use the Beam Search algorithm, introduce a diversity penalty term in the frequency division of each candidate sequence, and dynamically adjust the generation parameters according to the user's profile to generate a diverse dialogue sequence;
[0191] Take the said v fused as the input X sequence, X = [x1, x2, …, x n , and each x i is a word vector in the input sequence. The specific steps of the Transformer-based dialogue generation model are as follows:
[0192] See Figure 4 shown in the Transformer model encoder structure:
[0193] Position encoding: To solve the problem of loss of position information in the input sequence, add position encoding PE to each input vector to represent the position of each word in the input sequence:
[0194]
[0195] In the formula: t represents the position information of the word; i represents the vector dimension index; d represents the dimension of the vector;
[0196] Here, through the sin and cos functions, words in different positions have different encodings, avoiding the situation where the model cannot perceive the order in the sequence;
[0197] The position encoding PE(t) is added to the input word vector x i to obtain the enhanced input representation x' i = x i + PE(i);
[0198] Multi-head attention mechanism: Take the enhanced representation x'i It is input into the self-attention mechanism for processing. The self-attention mechanism calculates the relationship between each word and other words, that is, for the vector x' i to perform a linear transformation to obtain query, key, and value vectors:
[0199] Q t = X'W Q , K t = X'W K , V t = X'W V ;
[0200] In the formula: X' represents the input enhanced representation matrix (including positional encoding); W Q , W K and W V respectively represent trainable weight matrices;
[0201] Among them, in the self-attention mechanism, scaled dot-product attention is calculated:
[0202]
[0203] In the formula: Q represents the query matrix; K T represents the transpose of the key matrix; d k represents the dimension of the query vector or key vector; the correlation between each pair of words is calculated through dot product, then the weight of each word is calculated through the softmax function, and finally the weighted sum is obtained as the output;
[0204] In this embodiment, the multi-head attention mechanism is adopted. By calculating different attention heads in parallel, each head uses different weights of queries, keys, and values. Finally, the outputs of all heads are concatenated and passed through a linear transformation to obtain the final output, specifically as follows:
[0205] MultiHeat(Q, K, V) = Concat(head1, head2,..., head h )W O ;
[0206] In the formula: h represents the number of heads; W O represents the linear transformation matrix of the final output;
[0207] Feed-forward neural network: A feed-forward neural network is connected behind each attention layer. The feed-forward neural network includes two fully connected layers and an activation function (ReLU), and the calculation formula is as follows:
[0208] FFN(x) = max(0, xW1 + b1)W2 + b2;
[0209] where: W1 and W2 are trainable weight matrices; b1 and b2 represent bias terms;
[0210] Enhance the nonlinear ability of the model through two fully connected layers (linear transformation) and ReLU activation function;
[0211] Residual connection and layer normalization: Each sub-layer (self-attention layer and feed-forward neural network) is followed by a residual connection and layer normalization to ensure stable gradient transmission.
[0212] See Figure 5 as shown, the decoder structure of the Transformer model:
[0213] Input representation and positional encoding: The decoder input is first the word vector of the target sequence. Similar to the encoder, it first undergoes positional encoding processing:
[0214] y′ i = y i + PE(i);
[0215] where: y i represents the word vector of the target sequence; PE(i) represents the corresponding target positional encoding;
[0216] Causal self-attention mechanism: The first self-attention layer in the decoder uses causal self-attention to ensure that the generation of each word can only depend on the previous words and avoid future information leakage. The calculation of causal self-attention is the same as that of ordinary self-attention, but through masking technology, each position is restricted to only focus on itself and the previous positions;
[0217]
[0218] where: M represents the mask matrix, which makes future words have no impact on the calculation of the current word;
[0219] Cross-attention mechanism: The second attention layer in the decoder uses cross-attention, which correlates the encoder output with the current input of the decoder to help generate context-related outputs:
[0220]
[0221] where: Q comes from the input of the encoder, and K and V come from the output of the encoder;
[0222] Feed-forward neural network: The feed-forward neural network in the decoder is the same as that in the encoder;
[0223] Residual connection and layer normalization: Similar to the encoder, each sub-layer (self-attention layer and feed-forward neural network) in the decoder is followed by a residual connection and layer normalization;
[0224] Output layer: The final output of the decoder is generated through a linear transformation layer and a Softmax layer, mapping the output of the decoder to a probability distribution of the vocabulary size to select the next word to be generated;
[0225] The specific steps of adopting the Beam Search algorithm are as follows:
[0226] Generate candidate words: Based on the dialogue candidate sequence generated by the Transformer-based dialogue generation model, obtain the generation probability of each candidate word;
[0227] Diversity penalty: Calculate the diversity penalty for each candidate word:
[0228]
[0229] In the formula: represents the hidden state h of the current generated word y t ; t is the maximum cosine similarity between the hidden state h of the historical generated word y j ; λ represents the hyperparameter of the diversity penalty; j
[0230] Temperature adjustment: The generation probability of each candidate word is adjusted according to the user's style requirements. By performing temperature adjustment on the generation probability distribution, the randomness and diversity of the generation process are made to conform to the user's preferences:
[0231]
[0232] In the formula: T u represents the temperature adjustment factor;
[0233] Comprehensive scoring and ranking: The comprehensive score of each candidate sequence is adjusted based on the generation probability, historical similarity, and style requirements:
[0234]
[0235] In the formula: α represents the hyperparameter of the style adjustment factor; logP(y t |y1,y2…y t-1 ) represents the logarithmic probability when the candidate word y t is given the historical word sequence y1,y2…y t-1 ;
[0236] Based on the above steps, it is repeated at each time step. New candidate sequences are expanded according to the candidate sequence at the previous moment, and sorted according to the scores. Finally, the optimal candidate sequence is selected.
[0237] Step 7: Filter harmful content from the generated dialogue sequence through a rule layer, a model layer, and a post-processing layer to generate a safe dialogue sequence;
[0238] The rule layer is a rapid screening based on keyword and pattern matching, using regular expressions and keyword matching to quickly detect potential harmful content;
[0239] The model layer is a classification model based on the Transformer architecture, which classifies the content of each generated dialogue to determine whether the text contains inappropriate content;
[0240] The post-processing layer uses a BERT-based dialogue understanding model to analyze the rationality of the dialogue context, ensuring that the generated dialogue does not violate the previous dialogue background at the semantic level; and then, through the review and analysis of the historical dialogue, it judges whether the generated content is reasonable, and if it is not reasonable, manual intervention is carried out.
[0241] Currently, filtering harmful content using a rule layer, a model layer, and a post-processing layer belongs to the conventional technical means in this field, and will not be elaborated in detail here.
[0242] Step 8: Based on the safe dialogue sequence, construct a composite loss function based on cross-entropy loss, coherence loss, and style loss to optimize the dialogue generation model, obtaining an optimized dialogue generation model; generate dialogues based on the optimized dialogue generation model.
[0243] The composite loss function is as follows:
[0244]
[0245] In the formula: λ CE 、λ coh 、λ style represent the weight hyperparameters of cross-entropy loss, coherence loss, and style loss respectively; represents cross-entropy loss; represents coherence loss; represents style loss;
[0246]
[0247] In the formula: y i represents the true label distribution; represents the probability distribution predicted by the model; N represents the size of the vocabulary;
[0248] Among them, the coherence loss is measured by calculating the semantic distance between the currently generated sentence and the previous round of dialogue to measure its coherence:
[0249]
[0250] where: e(s t ) represents the semantic embedding representation of sentence s t ; e(s t-1 ) represents the semantic embedding representation of sentence s t-1 ; CosineSimilarity represents the cosine similarity function;
[0251] where the style loss is calculated based on the embedding vectors of user portrait features and sentiment analysis:
[0252]
[0253] where: e(u style ) represents the embedding of style-related features in the user portrait; represents the square of the Euclidean distance.
[0254] The present invention obtains authorized dialogue data through the WeChat Open Platform API, combines it with a public corpus, and through semantic anomaly detection using regular expressions and the BERT model, ensures that the cleaned dialogue dataset does not contain advertisements, spam, or meaningless dialogue fragments, thus guaranteeing the high quality and relevance of the data; secondly, when processing multi-turn dialogues, the dialogue data is organized by constructing a tree structure and introducing a three-dimensional timestamp, ensuring the effective processing of dialogue order and time information, thereby enhancing the coherence of the context; for the emotional information of emojis, the text and visual features of the emojis are combined through a hybrid encoder, enhancing the emotional expression in the dialogue and further improving the effect of personalized dialogue generation; in order to better model the dialogue context, a dynamic attention mechanism is adopted, which can weight according to the importance of historical sentences and adjust the memory length, thereby effectively retaining context information; after introducing user portrait features, by fusing user features with dialogue encoding, the generated dialogue becomes more personalized and meets the needs of different users; in the generation stage, through a Transformer-based model and the Beam Search algorithm, a diversity penalty term is introduced during the generation process and the generation parameters are dynamically adjusted according to the user portrait, thereby ensuring the diversity and security of the generated dialogue; finally, based on the optimization mechanism of the composite loss function, the accuracy and coherence of the dialogue generation model are further improved, effectively solving the technical problems of context coherence and personalized generation.
[0255] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for generating intelligent dialogues of digital pigeons on WeChat based on deep learning, characterized in that: The following steps are involved: Step 1: Obtain authorized conversation data based on the WeChat Open Platform API, and use the public conversation corpus as a supplementary data source; filter advertisements and spam in the conversation data and supplementary data sources based on regular expressions; and perform semantic anomaly detection through the pre-trained BERT model to identify and filter out meaningless or abnormal conversation fragments to obtain a cleaned conversation dataset; Step 2: Based on the cleaned conversation dataset, a tree structure is constructed to organize multiple rounds of conversations. Each conversation is organized into a tree structure based on timestamps and conversation contexts. The time information and conversation order are processed by attaching a three-dimensional timestamp to each message to obtain a reconstructed multi-round conversation dataset. Step 3: Construct corresponding sentiment values and text representations for the emojis in the multi-turn dialogue dataset, add sentiment information to the emojis by analyzing the sentiment values of the emojis, combine the text embedding and visual features of the emojis based on the hybrid encoder, and enhance the semantic representation of the emojis by weighting the text embedding and visual features, thus obtaining the multi-turn dialogue dataset after emoji processing; Step 4: Based on the multi-round dialogue dataset after expression matching processing, a dynamic attention mechanism is used to encode each historical sentence, the influence of each historical sentence is calculated, the historical sentences are weighted by the importance scoring model, and the memory length is adjusted according to the dialogue rounds to obtain the encoded context data; Step 5: Based on the encoded context data, user portrait features are introduced and mapped to the same vector space as the conversation encoding. The user features are mapped to a representation compatible with the conversation context through linear transformation to obtain a conversation dataset fused with user portraits. Step 6: Based on the dialogue dataset fused with user portraits, a Transformer-based dialogue generation model is used to generate dialogue candidate sequences. Based on the generated dialogue candidate sequences, the Beam Search algorithm is used to introduce a diversity penalty term in the frequency distribution of each candidate sequence. Based on the user portrait, the generation parameters are dynamically adjusted to generate diversified dialogue sequences. Step 7: Filter harmful content through the rule layer, model layer and post-processing layer to generate a safe dialogue sequence; Step 8: Based on the safe dialogue sequence, construct a composite loss function based on cross entropy loss, coherence loss, and style loss to optimize the dialogue generation model to obtain the optimized dialogue generation model; generate dialogue based on the optimized dialogue generation model.
2. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The steps for generating the reconstructed multi-round dialogue dataset are as follows: Time coding: Attach a three-dimensional timestamp to each message, including absolute timestamp, relative time difference, and order within the session; Time series embedding fusion: Based on the pre-trained BERT model, each message is embedded into text to obtain a text representation, and then the time series representation is obtained by embedding the timestamp, relative time difference and order within the conversation. The time series representation and text embedding are weightedly fused to obtain a reconstructed multi-round dialogue dataset.
3. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The step 3 comprises the following steps: Step 3.1: Build a mapping table containing expression symbols and emotion values, and define the emotion value for each expression symbol; Step 3.2: Based on the pre-trained BERT model, convert the text description of the expression into a high-dimensional vector representation to obtain the text embedding vector; Step 3.3: Use the pre-selected ResNet to process the image with the expression and obtain the visual feature vector; Step 3.4: The text embedding vector and the visual feature vector are fused by weighted averaging to obtain the expression representation vector; the weights in the weighted averaging are dynamically adjusted based on the conversation context; Step 3.5: Use the bidirectional LSTM model to modify the emotional value of the expression according to the context information, obtain the adjusted expression representation vector, and then obtain a multi-round dialogue data set after expression processing.
4. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The step 4 comprises the following steps: Step 4.1: Use a sliding window mechanism to retain the scheduled conversation turns; Step 4.2: In the determined dialogue round, based on the vector representation of each dialogue, the importance scoring model is used to calculate the attention weight: Where: α i Indicates historical dialogue h i The attention weight W q represents the query weight matrix; W k represents the key weight matrix; b represents the bias term; σ represents the Sigmoid activation function; R i Indicates historical dialogue h i An importance score of the content, wherein the importance score is calculated based on a time factor, an emotional factor, and a semantic relevance; Step 4.3: Use the calculated attention weights to weight the corresponding historical dialogues to obtain the weighted dialogue representation, i.e., the encoded context data.
5. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The step 5 comprises the following steps: Step 5.1: Extract user features, including interest tags, behavior data, preference information, and social attributes, and vectorize the extracted user features; Step 5.2: Take the weighted average of the user's interest tags, behavior data, preference information, and social attribute vectors to obtain a comprehensive user portrait vector; Step 5.3: Use linear transformation to map the user profile into a vector space that is the same as the encoded context data to obtain a projected user profile vector; Step 5.4: Use weighted fusion to fuse the encoded context data and the projected user profile; obtain a conversation data set that fuses the user profile, where the weights of the weighted fusion are dynamically adjusted based on learning of the neural network of the conversation content and user characteristics.
6. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The specific steps of using the Beam Search algorithm are as follows: Generate candidate words: Based on the dialogue candidate sequence generated by the Transformer dialogue generation model, obtain the generation probability of each candidate word; Diversity Penalty: Calculate the diversity penalty for each candidate word: Where: Indicates the current generated word y t The hidden state h t With the historical word y j The hidden state h j The maximum pre-similarity between ; λ represents the hyperparameter of diversity penalty; Temperature adjustment: The generation probability of each candidate word is adjusted according to the user's style requirements. By adjusting the temperature of the generation probability distribution, the randomness and diversity of the generation process are in line with the user's preferences: Where: T u represents the temperature adjustment factor; Comprehensive scoring and ranking: The comprehensive score of each candidate sequence is adjusted based on generation probability, historical similarity, and style requirements: Where: α represents the hyperparameter of the style adjustment factor; logP(y t |y1,y2…y t-1 ) indicates that when the candidate word y t Given a historical word sequence y1,y2…y t-1 The logarithmic probability of Based on the above steps, the process is repeated in each time step, a new candidate sequence is expanded based on the candidate sequence at the previous moment, and the new candidate sequence is sorted based on the scores, and finally the best candidate sequence is selected.
7. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The rule layer is a fast screening based on keyword and pattern matching, which uses regular expressions and keyword matching to quickly detect potential harmful content; The model layer is a classification model based on the Transformer architecture, which classifies the content of each generated conversation and determines whether the text contains inappropriate content; The post-processing layer uses the BERT-based dialogue understanding model to analyze the rationality of the dialogue context to ensure that the generated dialogue does not violate the previous dialogue context at the semantic level; Then, by reviewing and analyzing historical conversations, we determine whether the generated content is reasonable. If it is unreasonable, manual intervention will be carried out.
8. The method for generating a digital pigeon WeChat intelligent dialogue based on deep learning according to claim 1 is characterized in that: The composite loss function is as follows: Where: CE , coh , style Represent the weight hyperparameters of cross entropy loss, coherence loss and style loss respectively; represents the cross entropy loss; Indicates loss of coherence; represents style loss; Where: y i represents the true label distribution; represents the probability distribution predicted by the model; N represents the size of the vocabulary; The coherence loss measures the coherence of the currently generated sentence by calculating the semantic distance between it and the previous round of dialogue: Where: e(s t ) represents sentence s t The semantic embedding representation of e(s t-1 ) represents sentence s t-1 The semantic embedding representation of ; CosineSimilarity represents the cosine similarity function; The style loss is calculated based on the embedding vector of user portrait features and sentiment analysis: Where: e(u style ) represents the embedding of style-related features in user portraits; Represents the square of the Euclidean distance.
Citation Information
Patent Citations
Multi-round emotional dialogue method based on deep learning
CN108874972A
Personalized dialogue generation method and system based on user dialogue history
CN113779224A
Context deconstruction-based dialogue state tracking method and system
CN119272849A
Personalized dialogue content generating method
WO2021077974A1
Cited By
AI digital employee dialogue generation method and system based on deep learning
CN120542441A
AI intelligent toy multi-wheel automatic dialogue method based on large model
CN121071091A
Processing method and device for intelligent short message generation task
CN122242472A